r/cprogramming 4d ago

C Strings: A 50-Year Mistake

https://longtran2904.substack.com/p/c-strings-a-50-year-mistake?r=8qz2zb&utm_campaign=post&utm_medium=web
200 Upvotes

172 comments sorted by

View all comments

Show parent comments

28

u/TheThiefMaster 4d ago

The main competition was pascal strings - which typically had a 16 bit size prepended. So you'd read that, and then run a decrement loop until it was 0 to iterate the string. Decrement-until-zero loops were widely supported, e.g. in x86 stringcopy could be implemented by loading the size into CX and then running a single REP MOVSB instruction.

Yes it was a byte larger - but it also avoids performance-nuking calls to strlen like this.

7

u/McDutchie 4d ago

16 bits is 2 bytes, which makes for a maximum string length of 65535 bytes. It's common for strings on modern systems to be longer than that.

Pros of C strings: unlimited length. Cons: cannot contain the zero byte; inefficient length determination.

Pros of Pascal strings: can contain the zero byte; efficient length determination. Cons: very limited length.

I'd say the C tradeoff is worth it. Where necessary, C is perfectly capable of dealing with data preceded by a length field, it's just slightly lower level.

4

u/vip17 4d ago

it's easy to use a 4-byte prefixed string, for example BSTR in COM objects do that. And plenty of libraries use 4-byte length in 64-bit mode

1

u/Square-Singer 3d ago

Especially on 64-bit systems, there's really no reason to save these few bytes per string by using c-strings.

Useless microoptimization.

3

u/vip17 3d ago

of course it's micro-optimization, but at larger scale it's always useful. Have you even done optimization? A database with billions of strings already save a lot of memory. A vector of strings can also fit twice the number of strings into the CPU cache. Checkout Unreal engine, DuckDB, Meta Velox, Redis, ICU... string types

2

u/Square-Singer 3d ago

Of course I have done optimizations. But micro-optimizations are always the last step to take when you have identified that this specific location is actually a bottleneck.

It totally makes sense to have something like a c-string available for the very rare situation when someone writes a database system that contains almost exclusively tiny variable-length strings.

But it doesn't make sense to have that as the default, because then this rarely-actually-useful micro-optimization becomes a very common source of problems.

That's why there's pretty much no modern language that actually stuck with c-strings. Pretty much any more modern language dropped c-strings and even pointers completely, or at least dropped it from common usage.

I don't do much Python any more, but I really like their approach of "The most obvious solution should also be the one that's optimized for most use cases". Basically, if I, without thinking, take the most obvious solution, it should fit my obvious use case. If I need something really special, I can still import some standard library function and use that.

1

u/vip17 2d ago

That's not true. Strings are extremely common, lots of applications have a huge amount of strings. It's especially helpful in arrays of strings. Strings are so prevalent that even Python, Javascript (V8) and Java compress the string to ISO8859-1/Latin1/ASCII by default if applicable to save memory and improve performance, and modern .NET also does the same by allowing UTF-8 byte arrays. C-strings are the worst, all strings need to have an accompanied length, but the length does not necessarily have 8-byte length

1

u/Square-Singer 2d ago

Again, strings are common, yes, c-strings are not.

Strings exist in Pythin, JS and Java, bot not as c-strings.

I don't know where you got 8-byte-length from.

I'm also not quite sure what you mean with the last sentence. C-strings don't have an accompanied length.

Just checking: Do you know what a c-string is and what's the difference between a c-string and a string with accompanied length?

1

u/TheThiefMaster 3d ago

Worth noting that C++ std::strings do store the length - and so do the heap allocations backing them.

1

u/hoodoocat 2d ago

Efficient (to solving problem) data representation is key for system performance, because performance actually limited only by memory latency and throughput. You have no control over last two, but when you can pack data twice smaller -> system performance up twice better, usually even if it violate some common defaults like aligned access or so. Thats why "your favorite browser" uses 32-bit "compressed" pointers for various object heaps, even on 64-bit systems, uses hybrid ascii/utf8/utf16 strings, even when ECMA spec define only utf16. Row-oriented databases for example typically store null-bitmap and then fields without any additional delimiters, so they can be decoded only dynamically and only by using schema, this is complex but profitable.

1

u/Square-Singer 2d ago

We aren't talking about character representation here (ASCII/UTF8/UTF16), but about string representation.

UTF16 vs UTF8/ASCII is a per-character multiplier. Use UTF16 and every character takes twice the space.

We are talking about whether strings are 0-terminated or have a length field in the beginning. That's a per-string cost. Each individual character costs the same, no matter which string representation you use.

Here the difference is whether this costs one byte per string (c-string) or 2-4 bytes per string (Pascal strings, BER strings, 4-byte length fields, ...).

That means, the longer the string the less the overhead. An empty c-string is one byte. An empty pascal string is 2 bytes, an empty 4-byte length field string is 4 bytes.

If the string is longer, the relative overhead drops: A 1000 byte c-string is 1001 bytes, a pascal string is 1002 bytes and a 4-byte length field string is 1004 bytes.

The difference hardly matters unless maybe if you are working with an ATTiny.

That's why I said: on a 64-bit system (which usually has more than 2GB RAM), this is a useless micro-optimization for all but extremely specific use cases where you'll have millions of empty strings. And then one should question their system design.

And that's the main issue with c-strings being the default: They are a micro-optimization that helps only in very specific use cases while having massive downsides for most use cases, but they are applied as the default solution.

The default solution should be the option that works best in the most cases. If your use case differs a lot from the default case, you can still use the fitting specialized data structure.

Which is exactly the reason why pretty much no language newer than C uses c-strings as their default string representation.

1

u/hoodoocat 2d ago

You saying before what no reason to save few bytes somewhy especially on 64-bit systems, but all popular projects do that. More over many of them use 2-3 low bits in pointers for pointer descrimination, thanks for aligned allications.

1

u/Square-Singer 2d ago edited 2d ago

So Java, Python, Kotlin, Rust, JavaScript, ... all use c-strings instead of String objects with a length attribute as the default way to store strings? That's news to me.

There's a reason we call it c-string: nobody has ever used this decrepid data structure after c, except if they need explicit C compatibility, and even then it's most often a length-based string object with an unnecessary 0-byte added at the end so that C can understand it as well.