Repository navigation
Weird error using IO.copy_stream, IO duck types and enumerators #4903
Description
Activity
How peculiar! I'll have a look.
Really? I'd expect something else, I can reproduce it very consistently. How big was your image?
There's definitely something strange going on. With some larger input I am seeing something similar to what you reported, but it doesn't make any sense. The chunk produced by
@drain_stream.nextis passed back to the loop. It seems like something is continuing to modify that string, or it is malformed to begin with.yes, when I mean "consistently", I mean "repeating task and seeing it fail quite often". I'm not sure what I should attribute this to, but I'd say it is some internal buffering issue. My problem is "scaling down" my example to know where this actually happens, as if I try to simplify the code path, suddenly everything works flawlessly.
I still do not have an explanation for this. Poking at it a bit this afternoon.
Ok, I think I've figured out the issue.
Our Enumerator#next is frequently (usually) backed by a thread, since we do not have a way to do coroutines (e.g. lightweight Fibers) on the JVM. The logic for this thread should pause each time through the loop, waiting for the next item to be requested. In actuality, I believe it is immediately continuing to the next loop, which in this case results in the returned buffer getting overwritten before it can even be dup'ed.
I am trying to confirm this by examining logic in
copy_streamto see if it's reusing the same buffer repeatedly without marking it as shared.- added a commit that references this issue
on Jan 24, 2018 I've pushed a fix for the issue causing the data to be overwritten after it is returned. The triggering issue, Enumerator#next not waiting for a subsequent call to continue iterating, is in #5007 and will probably be fixed in next major release.
nice, thx for the fix! I'll test it as soon as I can.
@headius just tested this, and I can confirm the fix. Thx again for the top-level support!
- added 2 commits that reference this issue
on Feb 12, 2021
Environment
http-form_data, and with a jpg image (preferably with 46K)Expected Behavior
I have a very similar code to the one from this sample:
(The
putscalls are to debug and show the error)The purpose of this code is to enumerate the
IO.copy_streamcall, so that its chunks can be managed inside the while block. This code works in MRI (tested with 2.4).The main difference in implementation is that in MRI,
IO.copy_streamyields chunks of 16384 bytes, while JRuby yields 8192 bytes. I've followed this into this ticket, which leads me to believe that I can't reproduce this bug in older versions of jruby (as they were buffering the source in memory).If you limit the debug statements to
1., you'll see these outputs.In the end, the total bytes yielded in both solutions is similar. The gist of it is, the drained chunk must be equal to the yielded chunk.
However, if you limit the debug statements to
2., you'll see that this is not the case in JRuby:(check the 3rd yield)
Actual Behavior
As stated, I expect the pairs to be the same all the time.
I couldn't single out exactly what is the problem (The File buffer, the IO.copy_stream call, the enumeration...), and had to completely reproduce my usage to create this short script. But it's definitely a bug.