From 9296ec14e9efe901b397cf8560f75cacf867f170 Mon Sep 17 00:00:00 2001 From: Andrew Kelley Date: Thu, 20 Aug 2026 17:33:16 -0700 Subject: [PATCH] langref: rewrite Source Encoding section --- doc/langref.html.in | 53 +++++++++++++++------------------------------ 1 file changed, 17 insertions(+), 36 deletions(-) diff --git a/doc/langref.html.in b/doc/langref.html.in index 974acb49716d37308f88150d92c2a6ac3268127c..a0ea833d85a0f16320c71a784a4fd35a8a855f1a 100644 --- a/doc/langref.html.in +++ b/doc/langref.html.in @@ -7274,46 +7274,27 @@ fn readU32Be() u32 {} {#header_close#} {#header_close#} + {#header_open|Source Encoding#} -

Zig source code is encoded in UTF-8. An invalid UTF-8 byte sequence results in a compile error.

-

Throughout all zig source code (including in comments), some code points are never allowed:

+

Zig source code is UTF-8 encoded. Invalid UTF-8 byte sequences are not allowed anywhere.

+

Some code points are never allowed, even in {#link|Comments#}:

- LF (byte value 0x0a, code point U+000a, {#syntax#}'\n'{#endsyntax#}) is the line terminator in Zig source code. - This byte value terminates every line of zig source code except the last line of the file. - It is recommended that non-empty source files end with an empty line, which means the last byte would be 0x0a (LF). -

-

- Each LF may be immediately preceded by a single CR (byte value 0x0d, code point U+000d, {#syntax#}'\r'{#endsyntax#}) - to form a Windows style line ending, but this is discouraged. Note that in multiline strings, CRLF sequences will - be encoded as LF when compiled into a zig program. - A CR in any other context is not allowed. -

-

- HT hard tabs (byte value 0x09, code point U+0009, {#syntax#}'\t'{#endsyntax#}) are interchangeable with - SP spaces (byte value 0x20, code point U+0020, {#syntax#}' '{#endsyntax#}) as a token separator, - but use of hard tabs is discouraged. See {#link|Grammar#}. -

-

- For compatibility with other tools, the compiler ignores a UTF-8-encoded byte order mark (U+FEFF) - if it is the first Unicode code point in the source text. A byte order mark is not allowed anywhere else in the source. -

-

- Note that running zig fmt on a source file will implement all recommendations mentioned here. -

-

- Note that a tool reading Zig source code can make assumptions if the source code is assumed to be correct Zig code. - For example, when identifying the ends of lines, a tool can use a naive search such as /\n/, - or an advanced - search such as /\r\n?|[\n\u0085\u2028\u2029]/, and in either case line endings will be correctly identified. - For another example, when identifying the whitespace before the first token on a line, - a tool can either use a naive search such as /[ \t]/, - or an advanced search such as /\s/, - and in either case whitespace will be correctly identified. -

+ LF (byte value 0x0a, code point U+000a, {#syntax#}'\n'{#endsyntax#}) is + the line terminator in Zig source code. This byte value terminates every + line of Zig source code, including last line of the file. +

+

These conservative rules mean that third party tools reading + already-validated Zig source code may make simplifying assumptions, such + as naively separating lines based on {#syntax#}'\n'{#endsyntax#}. + However, tooling such as zig fmt provides convenience + functionality to convert invalid source encodings to valid source + encodings, for instance by stripping byte order marks and carriage + returns.

{#header_close#} {#header_open|Keyword Reference#} -- 2.54.0