Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.arch.embedded > #31433
| From | antispam@math.uni.wroc.pl |
|---|---|
| Newsgroups | comp.arch.embedded |
| Subject | Re: Serial Bus Speed on PCs |
| Date | 2022-12-05 03:33 +0000 |
| Organization | Aioe.org NNTP Server |
| Message-ID | <tmjoq2$1v7q$1@gioia.aioe.org> (permalink) |
| References | <1bf6f488-e849-4cfb-8625-ef968030cfban@googlegroups.com> <tm8uq3$1ac3$1@gioia.aioe.org> <ffc8758a-7a42-4e66-82a3-19e02d154866n@googlegroups.com> <tmj3hb$12h7$1@gioia.aioe.org> <36f1fd86-de78-49c1-bd31-bb1f35ad9ecen@googlegroups.com> |
Rick C <gnuarm.deletethisbit@gmail.com> wrote:
> On Sunday, December 4, 2022 at 4:30:35 PM UTC-5, anti...@math.uni.wroc.pl wrote:
> > Rick C <gnuarm.del...@gmail.com> wrote:
> > > On Wednesday, November 30, 2022 at 9:08:25 PM UTC-4, anti...@math.uni.wroc.pl wrote:
> > > > Rick C <gnuarm.del...@gmail.com> wrote:
> > One question is
> > if hardware is able to opperate with low latency. Another is if it
> > should. And frequently answer to secend question is no, it should
> > not try to minimize latency. Namely, Ethernet has minimal packet
> > size which is about 60 characters. If you send each character in
> > separate packet, then there would be very bad utilization of media.
> > So, converter is expected to wait till there is enough characters
> > to transmit.
>
> At the high serial rates we are talking about (3 Mbps) waiting for milliseconds is a bit absurd. Why? Because it can cause exactly the poor results I'm talking about. But the vendor with the Ethernet box never said the delay was intentional. What you are talking about would only be on the slave to master path. Why would data received from the network be delayed before sending to the slave? That sounds like a recipe for disaster.
Well, delay from Ethernet to serial port clearly means that implementer
did not spent enough effort to make it fast.
> > Note that at 115200 bits/s delay of 1ms is roughly
> > 11 characters, so not so big.
>
> We are not talking about 0.1 Mbps, rather 30x faster, 3 Mbps.
>
>
> > At lower rates delay becomes less
> > signifincant and at higher rates people usually care more about
> > throughput than latency. And do not forget that Ethernet is
> > shared medium, even if convertor could manage to transmit with
> > lower latency withing available Ethernet bandwidth, it could
> > do that only at cost of other users (possibly second convertor).
>
> Most Ethernet is not shared, rather point to point. In this case it definitely is not.
You were talking about connecting more convertors. Normally laptops
have only single Ethernet port, so all convertors that you connect
will share single Ethernet. If you use 100 convertors, 3 Mbits/s each
+ switches it should be possible to get 30 Mbytes/s of aggregate bandwidth
(assuming gigabyte port in laptop and gigabyte switch at top of tree).
But if each converter would waste a lot of bandwidth due to small
payload per packet, then such rate would be impossible.
> > And from a bit different point of view: normally there will be
> > software in the path, giving you 0.1ms of latency on good modern
> > unloaded hardware and much more in worse conditions.
>
> Ok, now it sounds like you are agreeing with me that the hardware is poor.
>
>
> > Also,
> > Ethernet likes packets of about 1400 bytes size. On 10 Mbit/s
> > Ethernet this is about 1.4 ms for transmitssion of packet.
>
> No one is using 10 Mbps. If needed, I would use 1 Gbps. I'm assuming that's not required. But you keep talking about "likes" and shared which don't apply here, at all. If a product is going to support up to 16 ports at 3 Mbps each, it seems like a bad idea to saddle them with such throughput killers.
It is your planned use that would kill throughput. I would expect
that when product is used as intended you would get resonable fraction
(say 70%) of nominal throughput (that is 2*16*3Mbits/s). If not,
then I will join you in calling it bad product.
> > If network in not dedicated to convertor such packets are likely
> > to appear from time to time and convertor has to wait till
> > such packet is fully transmitted and only then gets chance
> > to transmit. So, you should regularly expect delays of order
> > 1ms. Of course, with 100 Mbit/s Ethernet or gigabit one
> > media delays are smaller, but serial convertors are frequently
> > deployed in legacy contexts where 10 Mbit/s matter.
>
> Now you are being silly. If they design the equipment to work on 100 Mbps or even 1 Gbps Ethernet, you think it's reasonable for them to limit it to what they can do on 10 Mbps?
Sometimes you get product designed for 10 Mbit/s which just got faster
Ethernet part to be good citizen on fast network. Above you wrote
about 16 port thing. That should be designed for faster network, but
on common 100 Mbit/s Ethernet running ports in parallel it would be
limited by Ethernet troughput. And even on 1 Gbit/s Ethernet it
needs enough bandwidth that you can not waste it even if it is
the only thing on the network. And just a litte thing: you
wrote Ethernet, but raw Ethernet is problematic on PC OSes.
So I would guess that you really mean TCP/IP over Ethernet.
TCP requires every packet to be acknowleged, which may add more
small-packet trafic.
> > > > With relatively cheap convertors
> > > > on Linux to handle 10000 roundtrips for 15 bytes messages I need
> > > > the following times:
> > > >
> > > > CH340 2Mb/s, waiting, 6.890s
> > >
> > > That's 11.3 per target, per second. (128 targets)
> > >
> > > > CH340 2Mb/s, overlapped 1.058s
> > >
> > > That's pretty close to 74 per target, per second.
> > >
> > > I used to use the CH340 devices, but we had intermittent lockups of the serial port when testing all day long. I switched to FTDI and that went away. I think you told me you have no such problems. Maybe it's the CH340 Windows serial drivers.
> > Well, my use is rather light. Most is for debugging at say 9600 or
> > 115200. And when plugged in convertor mostly sits idle. I previously
> > wrote that CH340 did not work at 921600. More testing showed that
> > it actually worked, but speed was significantly different, I had to
> > set my MCU to 847000 communicate. This could be bug in Linux driver
> > (there is rather funky formula connecting speed to parameters
> > and it looks easy to get it wrong). Similary, when CH340 was set to 576800
> > I had to set MCU to 541300. Even after matching speed at nomial
> > 576800, 921600 and 1152000 test time was much (more than 10 times)
> > higher than for other rates (I only tested 1 character messages at those
> > rates, did not want to wait for full test). Also, 500000 was significantly
> > slower than 460800 (but "merely" 2 times slower for 1 character messages
> > and catching up with longer messages). Still, ATM CH340 looks
> > resonably good.
>
> Yes, it's reasonably good for situations where it does not need to work reliably. I was surprised when the finger was pointed to the CH340 adapter. But someone (probably here) had warned me they are not dependable, and now I know. The cost of a name brand adapter is not so much that it's worth saving the difference, only to have to throw it out and go with FTDI anyway, when you have real work to do.
Well, I say you what I observed. People say various thing on the
net. I was interested if net know something about my trouble with
CP2104 so I googled for "CP2104 lockup". And I got a bunch of
complaints about FTDI devices, solved by using CP2104. So, there
is a lot of noise and ATM I prefer to stay with what I see.
> > Remark: I bought all my convertors from Chinese sellers. IIUC
> > FTDI chip is faked a lot, but other too. Still, I think they
> > show what is possible and illustrate some difficulties.
>
> FTDI fakes no longer work with the FTDI drivers. Maybe they play a cat and mouse game, with each side one upping the other, but it's not worth the bother to try it out. FTDI sells cables. It's easier to just buy them from FTDI.
AFAIK Linux driver does not discriminate againt non-FTDI devices.
So fact that convertors works with Linux driver tells you nothing
about its origin. And for the record, I bought mine several years
ago.
> > > > CP2104 2Mb/s, waiting, 2.514s
> > > > CP2104 2Mb/s, overlapped 1.214s
> > >
> > > I don't know what the CP2104 is.
> > It is a chip by Silicon Laboratories. Datasheet gives contact address
> > in Austin, TX.
> > > I'm not certain what "overlapped" means in this test. Did you just continue to send 15 byte messages with no delays 10,000 times?
> > No. My slave simply returns back each received character. There is
> > some software delay but it should be less than 2us. So even waiting
> > test has some overlap at character level. To get more overlap above
> > I cheated: my test program was sending 1 more character than it should.
> > So sent message was 16 bytes, read was 15. After reading 15 another
> > batch of 16 was sent and so on. In total there were 10000 more
> > characters sent than received. My hope was that OS would read
> > and buffer excess characters, but it seems that at least for
> > CP2104 they cause trouble. My current guess is that OS is
> > reading only when requested, but I did not investigate deeper...
> > > Since you are in the mood for testing, what happens if you run overlapped, with 128 messages of 15 characters and wait for the replies before sending the next batch? Also, if you don't mind, can you try 20 character messages?
> > OK, I tried modifeed version of my test program. It first sends
> > k messages without reading anything, then goes to main loop where
> > after sending each message it read one. At the end it tail loop
> > which reads last k messages without sending anything. So, there
> > is k + 1 messages in transit: after sending message k + i program
> > waits for answer to message i. In total there is 10000 messages.
> > Results are:
> >
> > CH340, 15 char message 20 char message
> > k = 0 6.869s 7.163s
> > k = 1 4.682s 1.320s
> > k = 2 0.992s 1.320s
> > k = 3 0.991s 1.319s
> > k = 4 0.991s 1.320s
> > k = 5 0.990s 1.319s
> > k = 8 0.992s 1.320s
> > k = 12 0.990s 1.320s
> > k = 20 0.992s 1.319s
> > k = 36 0.991s 1.321s
> > k = 128 0.991s 1.319s
> >
> > CP2104, 15 char message 20 char message
> > k = 0 2.508s 3.756s
> > k = 1 1.897s 1.993s
> > k = 2 1.668s 2.087s
> > k = 3 1.486s 1.887s
> > k = 4 1.457s 1.917s
> > k = 5 1.559s 1.877s
> > k = 8 1.455s 1.803s
> > k = 12 1.337s 1.501s
> > k = 20 1.123s 1.499s
> > k = 36 1.125s 1.502s
> >
> > k = 128 reliably stalled, there were random stalls in other cases
> >
> > FTDI232R,
> > 2 Mbit/s 15 char message 20 char message
> > k = 0 5.478s 3.755s
> > k = 1 4.929s 3.030s
> > k = 2 2.506s 3.339s
> > k = 3 2.459s 2.020s
> > k = 4 1.708s 1.061s
> > k = 5 1.671s 1.032s
> > k = 8 0.764s 1.021s
> > k = 12 0.772s 1.014s
> > k = 20 0.763s 1.009s
> > k = 36 0.758s 1.007s
> > k = 128 0.757s 1.008s
> >
> > FTDI232R,
> > 3 Mbit/s 15 char message 20 char message
> > k = 0 8.216s 10.007s
> > k = 1 5.006s 4.344s
> > k = 2 3.338s 1.602s
> > k = 3 2.406s 1.444s
> > k = 4 1.766s 1.316s
> > k = 5 1.599s 1.673s
> > k = 8 1.040s 1.327s
> > k = 12 1.071s 1.312s
> >
> > With k = 20, k = 36 and k = 128 communication stalled.
>
> Some of the results seem odd, hard to understand, like why the message rate improves so much as k is increased, but so dramatically at 3 Mbps. They all seem to approach ~1.3 second as k increases. At k=0 they are around 1 ms per message, which is the polling rate... if you adjust it. I think the default for FTDI was 8 ms.
Let me first comment 2Mbit/s results. FTDI transfers data in 64-byte
blocks (they say that actual payload is 62-bytes and there are 2-bytes
of protocol info). With 15 characters messages 0.764s really means
98% of use of serial bandwidth, so essentiall as good as possible.
Corresponding k = 8 means really 9 messages in transit, so 135
characters which is slightly more than 2 buffers. More data in
transit does not help, but also does not make things worse.
With 20 charaster messages main improvement is at k = 4 which
means 100 characters, which is smaller than 2 buffers, with extra
improvements for more data in transit. With CH340 and 15 char
messages we see main improvement for k = 2, which corresponds
to 45 characters in transit. With 20 char messages we get
impovement for k = 1 which is 40 charactes in transit.
CH340 uses 32 character transfer buffers, so improvemnet corresponds
to somwhat more than 1 buffer in transit. Now, if transfers
between converter and PC were at optimal times, then one buffer
+ one character would be enough to get full serial speed. But
USB tranfers can not be started at arbitrary times, IIUC there
are discrete time slots when transfer can occur. When tranfer
can not be done in given slot it must wait for next slot.
So, depending on locations of possible slots more buffering
and more data in transit may be needed for optimal performance.
OTOH 2-3 buffers should be enough to allow PC to get full
bandwidth and this is in good agreement with FTDI results.
In case of CH340 there is extra factor: CH340 also uses 8 byte
transfers. I do not know what function they have, but
resonably likely guess is that those 8 byte pack tranfer control
info that FTDI bundles with normal data. Anyway, those
are "interrupt" tranfers in USB sense, so have higher priority
than data transfer. Resonable guess it that they steal some
USB bandwith from data tranfers. Also, smaller than maximal
data block size limits efficiency, so it is possible that
CH340 is limited by USB bandwith (lack of enough slots).
Now, concerning 3 Mbits/s, due to different serial speed
optimal times for transfers are different than in 2 Mbits/s
case. It is possible that there is worse fit of desired
and possible transfer times. Buffering allows to at least
partially cure this, so initial improvement. But clearly,
there is some extra bottleneck. Now some speculation:
with 1/8 ms USB-2.0 cycle, there is 1500 FS clock per
cycle. I would have to look at spec to be sure, but this
is close to 150 byte worst case FS transfer. Beside data
there is some USB protocol overhead and (speculatively) it
is possible that low level USB diver may refuse to schedule
two 64-byte transfers in single cycle. In such case effective
bandwith for serial data would be 4096000 bits, which
correspond to 5120000 serial bits (serial sends start and stop
bits which are not needed for USB). This is less than
full duplex 3 Mbits/s (both directions add to 6 Mbits/s and
must go trouh the same USB). With larger amount of data in
transit this could give wild oscilations in amount of
buffered data, leading to slowdown when buffers get empty
and giving stall when receive buffer overflows.
Of course there is another speculation: convertor may be fake.
Supposedly fakes use MCU-s with special program. Software
could crate delays which limit transfer rate at 3 Mbits/s
and lead to data loss/stall with more data in transit.
> > > > As other suggested you could use multiple convertors for
> > > > better overlap. My convertors are "full speed" USB, that
> > > > is they are half-duplex 12 Mb/s. USB has significant
> > > > protocol overhead, so probably two 2 Mb/s duplex serial
> > > > convertes would saturate single USB bus. In desktops
> > > > it is normal to have several separate USB controllers
> > > > (buses), but that depends on specific motherboard.
> > > > Theoreticaly, when using "high speed" USB converters,
> > > > several could easily work from single USB port (provided
> > > > that you have enough places in hub(s)).
> > >
> > > I've been shying away from USB because of the inherent speed issues with small messages. But with larger messages, hi-speed converters can work, I would hope. Maybe FTDI did not understand my question, but they said even on the hi-speed version, their devices use a polling rate of 1 ms. They call it "latency", but since it is adjustable, I think it is the same thing. I asked about the C232HD-EDHSP-0, which is a hi-speed device, but also mentioned the USB-RS422-WE-5000-BT, which is an RS-422, full-speed device. So maybe he got confused. They don't offer many hi-speed devices.
> > >
> > > But the Ethernet implementations also have speed issues, likely because they are actually software based.
> > The issues are more fundamental: both in USB and Ethernet there
> > is per message/packet overhead. Low latency means sending data
> > soon after it is available, which means small packets/messages.
> > But due to overheads small packets are bad for throughput.
> > So designers have to choose what they value more and in both
> > cases the whole system is normally optimized for throughput.
>
> With 100 Mbps Ethernet the inherent latencies are very low compared to my message transmission rates. One vendor specifically indicated the delays were in their software. That was when I mentioned FPGAs and he talked as if I were being ridiculous.
Well, you wrote that you have needed experience, so do low-latency
Ethernet-serial convertor based on FPGA. Give your numbers and look
how many customers come in.
--
Waldek Hebisch
Back to comp.arch.embedded | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-11-29 23:33 -0800
Re: Serial Bus Speed on PCs Richard Damon <Richard@Damon-Family.org> - 2022-11-30 07:42 -0500
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-11-30 06:21 -0800
Re: Serial Bus Speed on PCs Bernd Linsel <bl1-removethis@gmx.com> - 2022-11-30 16:11 +0100
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-11-30 07:58 -0800
Re: Serial Bus Speed on PCs David Brown <david.brown@hesbynett.no> - 2022-11-30 18:14 +0100
Re: Serial Bus Speed on PCs Dimiter_Popoff <dp@tgi-sci.com> - 2022-11-30 20:52 +0200
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-03 12:42 -0800
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-03 12:55 -0800
Re: Serial Bus Speed on PCs David Brown <david.brown@hesbynett.no> - 2022-12-04 13:21 +0100
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-04 08:54 -0800
Re: Serial Bus Speed on PCs David Brown <david.brown@hesbynett.no> - 2022-12-05 08:57 +0100
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-05 14:12 -0800
Re: Serial Bus Speed on PCs antispam@math.uni.wroc.pl - 2022-12-01 01:08 +0000
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-01 02:48 -0800
Re: Serial Bus Speed on PCs David Brown <david.brown@hesbynett.no> - 2022-12-02 13:30 +0100
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-02 09:01 -0800
Re: Serial Bus Speed on PCs antispam@math.uni.wroc.pl - 2022-12-04 21:30 +0000
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-04 14:57 -0800
Re: Serial Bus Speed on PCs antispam@math.uni.wroc.pl - 2022-12-05 03:33 +0000
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-04 22:39 -0800
Re: Serial Bus Speed on PCs antispam@math.uni.wroc.pl - 2022-12-06 02:30 +0000
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-05 23:41 -0800
Re: Serial Bus Speed on PCs David Brown <david.brown@hesbynett.no> - 2022-12-06 13:32 +0100
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-06 18:03 -0800
Re: Serial Bus Speed on PCs David Brown <david.brown@hesbynett.no> - 2022-12-07 08:08 +0100
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-07 03:36 -0800
Re: Serial Bus Speed on PCs Andrew Smallshaw <andrews@sdf.org> - 2022-12-05 09:58 +0000
Re: Serial Bus Speed on PCs Rick C <gnuarm.deletethisbit@gmail.com> - 2022-12-05 14:03 -0800
csiph-web