Standard `Debug` implementation for `[u8]` is comma separated list
of numbers. Since large amount of byte strings are in fact ASCII
strings or contain a lot of ASCII strings (e. g. HTTP), it is
convenient to print strings as ASCII when possible.
Limit the number of threads when using qemu to 1. Also, don't bother
running the stress test as this will trigger qemu bugs. Finally, also
make the stress test actually stress test.
I found this significantly improved a
[benchmark](https://gist.github.com/danburkert/34a7d6680d97bc86dca7f396eb8d0abf)
which calls `bytes_mut`, writes 1 byte, and advances the pointer with
`advance_mut` in a pretty tight loop. In particular, it seems to be the
inline annotation on `bytes_mut` which had the most effect. I also took
the opportunity to simplify the bounds checking in advance_mut.
before:
```
test encode_varint_small ... bench: 540 ns/iter (+/- 85) = 1481 MB/s
```
after:
```
test encode_varint_small ... bench: 422 ns/iter (+/- 24) = 1895 MB/s
```
As you can see, the variance is also significantly improved.
Interestingly, I tried to change the last statement in `bytes_mut` from
```
&mut slice::from_raw_parts_mut(ptr, cap)[len..]
```
to
```
slice::from_raw_parts_mut(ptr.offset(len as isize), cap - len)
```
but, this caused a very measurable perf regression (almost completely
negating the gains from marking bytes_mut inline).
The `Source` trait was essentially covering the same case as `IntoBuf`,
so remove it.
While technically a breaking change, this should not have any impact due
to:
1) There are no reverse dependencies that currently depend on `bytes`
2) Source was not supposed to be implemented externally
3) IntoBuf provides the same implementations as `Source`
Given these points, the change should be safe to apply.
This change tracks the original capacity requested when `BytesMut` is
first created. This capacity is used when a `reserve` needs to allocate
due to the current view being too small. The newly allocated buffer will
be sized the same as the original allocation.
The previous implementation didn't factor in a single `Bytes` handle
being stored in an `Arc`. This new implementation correctly impelments
both `Bytes` and `BytesMut` such that both are `Sync`.
The rewrite also increases the number of bytes that can be stored
inline.