I found this significantly improved a
[benchmark](https://gist.github.com/danburkert/34a7d6680d97bc86dca7f396eb8d0abf)
which calls `bytes_mut`, writes 1 byte, and advances the pointer with
`advance_mut` in a pretty tight loop. In particular, it seems to be the
inline annotation on `bytes_mut` which had the most effect. I also took
the opportunity to simplify the bounds checking in advance_mut.
before:
```
test encode_varint_small ... bench: 540 ns/iter (+/- 85) = 1481 MB/s
```
after:
```
test encode_varint_small ... bench: 422 ns/iter (+/- 24) = 1895 MB/s
```
As you can see, the variance is also significantly improved.
Interestingly, I tried to change the last statement in `bytes_mut` from
```
&mut slice::from_raw_parts_mut(ptr, cap)[len..]
```
to
```
slice::from_raw_parts_mut(ptr.offset(len as isize), cap - len)
```
but, this caused a very measurable perf regression (almost completely
negating the gains from marking bytes_mut inline).
The `Source` trait was essentially covering the same case as `IntoBuf`,
so remove it.
While technically a breaking change, this should not have any impact due
to:
1) There are no reverse dependencies that currently depend on `bytes`
2) Source was not supposed to be implemented externally
3) IntoBuf provides the same implementations as `Source`
Given these points, the change should be safe to apply.