Skip to content

Replace malformed unicode error helper calls with the _real variants#8359

Open
devyubin wants to merge 1 commit into
RustPython:mainfrom
devyubin:unicode-error-real
Open

Replace malformed unicode error helper calls with the _real variants#8359
devyubin wants to merge 1 commit into
RustPython:mainfrom
devyubin:unicode-error-real

Conversation

@devyubin

@devyubin devyubin commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Part of #8352

Summary

new_unicode_decode_error / new_unicode_encode_error construct a structurally invalid exception: it has none of the five attributes a unicode error must carry (encoding, object, start, end, reason), its args is a 1-tuple message, and str() renders as an empty string. This converts the non-Windows call sites to the _real constructors:

import csv
next(csv.reader(['\udcff'], quoting=csv.QUOTE_STRINGS))
# before: str(e) = ''            args = ('csv not utf8',)   e.encoding -> AttributeError
# after:  str(e) = "'utf-8' codec can't decode byte 0xed in position 0: csv not utf8"
#         args = ('utf-8', b'\xed\xb3\xbf', 0, 1, 'csv not utf8')   all five attributes set

Cause

The broken helpers are generated by define_exception_fn!, which goes through new_exception_msgPyBaseException::new(args). That only stores args and never runs the type's initializer, so the unicode-error attributes are never set. The _real constructors pass the full (encoding, object, start, end, reason) and set every attribute.

Fix

Converted call sites, supplying the source bytes/str and the failing offset from the previously discarded Utf8Error (valid_up_to(), span start..start + 1 matching the existing _real usages):

  • csv: all 6 sites, via a new_not_utf8_error helper next to the existing new_csv_error
  • socket: host/port bytes decode; the port surrogate check now reports the first surrogate position (same scan as PyStr::ensure_valid_utf8)
  • posix.getlogin, FsPath::bytes_as_os_str, os::bytes_as_os_str

Left for follow-ups, per the issue plan:

  • the Windows-gated sites (nt.rs, mbcs/oem codecs) — I can't test Windows locally
  • sites whose source object is not reachable at the error site (uname, array) — these need a small refactor or an error-type decision rather than a mechanical conversion

Test

  • Repro above now produces a well-formed UnicodeDecodeError; also verified socket.getaddrinfo(b"\xff\xfe", 80)('utf-8', b'\xff\xfe', 0, 1, ...) and socket.getaddrinfo("localhost", "ab\udcffcd")UnicodeEncodeError ('utf-8', 'ab\udcffcd', 2, 3, 'surrogates not allowed') (surrogate position reported precisely).
  • cargo run -- -m test test_csv / test_os / test_posixTests result: SUCCESS (no unexpected successes).
  • cargo clippy / cargo fmt: clean (no new warnings).

Assisted-by: Claude Code:claude-fable-5

Summary by CodeRabbit

  • Bug Fixes
    • Improved Unicode decoding errors for invalid CSV data, including clearer encoding details and error locations.
    • Improved socket address handling errors for invalid host or port text and unsupported surrogate characters.
    • Improved error reporting when filesystem paths or login names contain invalid UTF-8.
    • Preserved existing CSV processing, DNS lookup, path handling, and login behavior for valid input.

new_unicode_decode_error / new_unicode_encode_error build the exception
from a bare message without running the initializer, so the result has
none of the five attributes a unicode error must carry and str() renders
as an empty string. Convert the non-Windows call sites in csv, socket,
getlogin and the fs-path decoders to the _real constructors, passing the
source bytes/str and the failing offset from the captured Utf8Error.

The Windows-gated sites (nt, mbcs/oem codecs) and the sites whose source
object is not reachable (uname, array) are left for follow-ups.

Assisted-by: Claude Code:claude-fable-5
@coderabbitai

coderabbitai Bot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

CSV, socket, filesystem, OS, and POSIX code now construct detailed Unicode errors with encoding, source bytes, and invalid-position information instead of generic messages.

Changes

Unicode Error Reporting

Layer / File(s) Summary
CSV UTF-8 error handling
crates/stdlib/src/csv.rs
Adds a shared helper and applies structured UTF-8 decode errors across CSV readers and writer output paths.
Socket address argument errors
crates/stdlib/src/socket.rs
getaddrinfo() now reports detailed host and port decoding failures and surrogate encoding errors.
Filesystem and login decoding errors
crates/vm/src/function/fspath.rs, crates/vm/src/stdlib/os.rs, crates/vm/src/stdlib/posix.rs
Path and login-name decoding failures now include source bytes and invalid UTF-8 positions.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Suggested reviewers: shaharnaveh, joshuamegnauth54

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: switching malformed Unicode error helpers to the _real variants.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/stdlib/src/csv.rs`:
- Around line 63-76: Update new_not_utf8_error to calculate the decode-error end
offset from err.error_len(): use err.valid_up_to() plus the reported length when
available, and bytes.len() for truncated sequences when it is None. Pass this
computed end offset to vm.new_unicode_decode_error_real instead of always using
err.valid_up_to() + 1, preserving the complete invalid UTF-8 span.

In `@crates/stdlib/src/socket.rs`:
- Around line 2615-2622: Update both host and port UTF-8 error handlers around
the host and port from_utf8 calls in crates/stdlib/src/socket.rs:2615-2622 and
crates/stdlib/src/socket.rs:2656-2664 to calculate the UnicodeDecodeError end
position using Utf8Error::error_len(), falling back to the remaining input
length when needed, instead of always using valid_up_to() + 1.

In `@crates/vm/src/function/fspath.rs`:
- Around line 128-136: Update the UTF-8 error mappings in
crates/vm/src/function/fspath.rs:128-136, crates/vm/src/stdlib/os.rs:133-141,
and crates/vm/src/stdlib/posix.rs:1734-1742 to use the complete failure span:
derive the end offset from e.error_len(), falling back to the source slice
length when it is None, and use the source length for unexpected EOF instead of
valid_up_to() + 1.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yml

Review profile: CHILL

Plan: Pro Plus

Run ID: 41bb1195-e3a9-4784-913d-1463dff8e1a0

📥 Commits

Reviewing files that changed from the base of the PR and between 28454cc and 02c70c4.

📒 Files selected for processing (5)
  • crates/stdlib/src/csv.rs
  • crates/stdlib/src/socket.rs
  • crates/vm/src/function/fspath.rs
  • crates/vm/src/stdlib/os.rs
  • crates/vm/src/stdlib/posix.rs

Comment thread crates/stdlib/src/csv.rs
Comment on lines +63 to +76
fn new_not_utf8_error(
vm: &VirtualMachine,
bytes: &[u8],
err: core::str::Utf8Error,
) -> PyBaseExceptionRef {
vm.new_unicode_decode_error_real(
vm.ctx.new_str("utf-8"),
vm.ctx.new_bytes(bytes.to_vec()),
err.valid_up_to(),
err.valid_up_to() + 1,
vm.ctx.new_str("csv not utf8"),
)
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail
rg -n "new_not_utf8_error|csv not utf8|UnicodeDecodeError" crates/stdlib crates/vm

Repository: RustPython/RustPython

Length of output: 1979


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== file outline =="
ast-grep outline crates/stdlib/src/csv.rs --match new_not_utf8_error --view expanded || true

echo "== relevant csv.rs section =="
sed -n '55,85p' crates/stdlib/src/csv.rs | cat -n

echo "== call sites context =="
for range in "1115,1135" "1258,1272" "1435,1447" "1480,1492" "1530,1542" "1604,1616"; do
  echo "--- range $range ---"
  sed -n "${range}p" crates/stdlib/src/csv.rs | cat -n -v
done

echo "== Python UnicodeDecodeError semantic probe =="
python3 - <<'PY'
samples = {
    "multi_byte_invalid": b"\xf0\xe2\x88\x92"
}
for name, data in samples.items():
    try:
        s = data.decode("utf-8")
    except UnicodeDecodeError as e:
        print(name, {
            "encoded": list(e.object),
            "start": e.start,
            "end": e.end,
            "span": list(e.object[e.start:e.end]),
            "reason": e.reason
        })
PY

echo "== Rust Utf8Error probe (no repo code execution, using rustdoc/std from local if available) =="
cat > /tmp/rust_utf8_error_probe.rs <<'RS'
use std::str;

fn main() {
    for (name, bytes) in [
        ("multi_byte_invalid", b"\xf0\xe2\x88\x92"),
        ("truncated_last_byte", b"abc\xf0\x9f"),
        ("single_byte_invalid", b"\x80"),
    ] {
        let e = str::from_utf8(bytes).unwrap_err();
        println!(
            "{} valid_up_to={} error_len={:?} span_from_valid_up_to_to_valid_up_to_plus_1_is=[]",
            name,
            e.valid_up_to(),
            e.error_len(),
            &bytes[e.valid_up_to()..e.valid_up_to()+1],
        );
        let start = e.valid_up_to();
        let end = e.error_len().map_or(bytes.len(), |len| start + len);
        println!(
            "  proposed span={:?} object_slice_start={:?} object_slice_end={:?}",
            &bytes[start..end],
            &bytes[start..start+1],
            &bytes[end-1..end],
        );
    }
}
RS
rustc /tmp/rust_utf8_error_probe.rs -o /tmp/rust_utf8_error_probe 2>/tmp/rustc_stderr.txt || { echo "rustc unavailable"; cat /tmp/rustc_stderr.txt; exit 0; }
/tmp/rust_utf8_error_probe

Repository: RustPython/RustPython

Length of output: 5740


🌐 Web query:

Rust core::str::Utf8Error error_len documentation valid_up_to truncated invalid_len None

💡 Result:

In Rust, the core::str::Utf8Error struct provides details about why a UTF-8 validation attempt failed [1][2]. Two primary methods used to inspect this error are valid_up_to and error_len [1][3]. valid_up_to The valid_up_to method returns the index (as a usize) up to which the input was confirmed to be valid UTF-8 [1][4]. It is defined as the maximum index such that calling from_utf8 on the slice up to that index would return Ok(_) [1][3]. error_len The error_len method returns an Option providing further information about the nature of the invalid sequence starting at the position returned by valid_up_to [1][5]. 1. Some(len): This indicates that an unexpected byte was encountered [1][6]. The value len represents the length of the invalid byte sequence [1][4]. In cases of lossy decoding, you should insert a U+FFFD REPLACEMENT CHARACTER and resume decoding after this invalid sequence [1][3]. 2. None: This indicates that the input ended unexpectedly [1][4]. The index returned by valid_up_to is between 1 and 3 bytes from the end of the input [1][6]. This often occurs when a multi-byte UTF-8 character is truncated because the input stream ended prematurely (e.g., when decoding a file or network socket incrementally) [1][4][3]. In this scenario, the bytes starting at valid_up_to might form a valid character if more data were available [1][6].

Citations:


Preserve the complete invalid UTF-8 span.

For Utf8Error, use error_len() where available; use bytes.len() for truncated sequences where error_len() returns None, so UnicodeDecodeError.end and the rendered message include the full invalid byte span.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/stdlib/src/csv.rs` around lines 63 - 76, Update new_not_utf8_error to
calculate the decode-error end offset from err.error_len(): use
err.valid_up_to() plus the reported length when available, and bytes.len() for
truncated sequences when it is None. Pass this computed end offset to
vm.new_unicode_decode_error_real instead of always using err.valid_up_to() + 1,
preserving the complete invalid UTF-8 span.

Comment on lines +2615 to +2622
let host_str = core::str::from_utf8(&bytes).map_err(|e| {
vm.new_unicode_decode_error_real(
vm.ctx.new_str("utf-8"),
vm.ctx.new_bytes(bytes.to_vec()),
e.valid_up_to(),
e.valid_up_to() + 1,
vm.ctx.new_str("host bytes is not utf8"),
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "Relevant lines:"
sed -n '2590,2675p' crates/stdlib/src/socket.rs | nl -ba -v2590

echo
echo "Search for Utf8Error/error_len/valid_up_to usage:"
rg -n "Utf8Error|error_len\\(|valid_up_to\\(|new_unicode_decode_error_real|from_utf8" crates/stdlib/src/socket.rs crates || true

echo
echo "Inspect rustc stdlib source docs locally if available:"
python3 - <<'PY'
import os, pathlib, subprocess, json
out = subprocess.run(["rustc","--version"], capture_output=True, text=True)
print("rustc:", out.stdout.strip())
root = subprocess.run(["rustc","--print","sysroot"], capture_output=True, text=True, check=True).stdout.strip()
for name in ["liballoc-*.rlib","libstd-*.rlib"]:
    print(root, name, len(list(pathlib.Path(root).glob(name))))
PY

echo
echo "Behavior probe from Rust Utf8Error without running repo code:"
python3 - <<'PY'
import subprocess
import tempfile, pathlib, textwrap, os, sys

src = r'''
use core::str::Utf8Error;

fn main() {
    const cases: &[&str] = &[
        r"abc\xFFdef",
        r"abc\xC0\x80def",
        r"abc\xC4\x80def",
        r"🐶abc",
        r"abc\xED\xF3\0def",
        r"abc",
    ];
    for case in cases {
        let bytes = case.as_bytes();
        println!("bytes: {:?}", String::from_utf8(c.clone()).err());
    }
}
'''
# no-op compile/check only; use exact small program patterns via valid_utf8 invalid pattern.
# Actually produce exact invalid UTF8 bytes probe:
src = r'''
fn main() {
    let cases: &[&[u8; 5]] = &[
        b"abc\xff",
        b"abc\xc0",
        b"abc\xc4",
        #[allow(clippy::len_without_is_empty)]
        &[][1..],
    ];
    println!("cases: {}", cases.len());
}
'''
tmp = tempfile.mkdtemp()
file = pathlib.Path(tmp) / "utf8probe.rs"
file.write_text(r'''
use core::str::Utf8Error;

fn main() {
    let cases: &[&[u8]] = &[
        b"abc\xff",
        b"abc\xc0",
        b"abc\xc4",
        &[0],
        b"abc\xed\xf3\x99\xbc",
        b"abc\xed\x9f\x80",
    ];
    for bytes in cases {
        let s = std::str::from_utf8(bytes);
        if let Err(e) = s.as_ref() {
            let start = e.valid_up_to();
            let hard = start + 1;
            let actual_hard_len = start.saturating_add(e.error_len().unwrap());
            use std::str::FromStr;
            let actual = core::str::error_len::core::str::Utf8Error(error_len)
            println!("case {:?}", bytes);
        }
    }
}
''')
print("compiled", subprocess.run(["rustc","--edition","2021","-q","/nonexistent"], capture_output=True, text=True).returncode)
print("not compiling probe since syntax intentionally invalid no-op")
PY

Repository: RustPython/RustPython

Length of output: 215


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "Relevant socket range:"
sed -n '2590,2675p' crates/stdlib/src/socket.rs | awk '{printf "%6d  %s\n", NR+2589, $0}'

echo
echo "Search related symbols/files:"
rg -n "Utf8Error|error_len\\(|valid_up_to\\(|new_unicode_decode_error_real|from_utf8" crates/stdlib/src/socket.rs crates 2>/dev/null | head -100

echo
echo "Utf8Error behavior probe:"
tmp="$(mktemp -d)" && chmod +w "$tmp"
cat > "$tmp/utf8probe.rs" <<'RS'
use core::str::Utf8Error;

fn main() {
    let cases: &[&[u8]] = &[
        b"abc\xff",
        b"abc\xc0",
        b"abc\xc4",
        &[0],
        b"abc\xed\xf3\x99\xbc",
        b"abc\xed\x9f\x80",
    ];
    for bytes in cases {
        let res: Result<&str, Utf8Error> = core::str::from_utf8(bytes);
        if let Err(e) = res {
            let start = e.valid_up_to();
            let hard_end = start + 1;
            let actual_len = e.error_len().unwrap_or(bytes.len() - start);
            let actual_end = start + actual_len;
            println!("bytes={:?} valid_up_to={} hardened_end={} actual_len={} actual_end={}", bytes, start, hard_end, actual_len, actual_end);
        } else {
            println!("bytes={:?} ok", bytes);
        }
    }
}
RS
rustc --edition 2021 "$tmp/utf8probe.rs" -o "$tmp/utf8probe" -q
"$tmp/utf8probe"

Repository: RustPython/RustPython

Length of output: 14706


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "Utf8Error behavior probe:"
tmp="$(mktemp -d)"
cat > "$tmp/utf8probe.rs" <<'RS'
use core::str::Utf8Error;

fn main() {
    let cases: &[&[u8]] = &[
        b"abc\xff",
        b"abc\xc0",
        b"abc\xc4",
        &[0],
        b"abc\xed\xf3\x99\xbc",
        b"abc\xed\x9f\x80",
    ];
    for bytes in cases {
        let res: Result<&str, Utf8Error> = core::str::from_utf8(bytes);
        if let Err(e) = res {
            let start = e.valid_up_to();
            let hard_end = start + 1;
            let actual_len = e.error_len().unwrap_or(bytes.len() - start);
            let actual_end = start + actual_len;
            println!("bytes={:?} valid_up_to={} hard_end={} actual_len={} actual_end={}", bytes, start, hard_end, actual_len, actual_end);
        } else {
            println!("bytes={:?} ok", bytes);
        }
    }
}
RS
rustc --edition 2021 -q "$tmp/utf8probe.rs" -o "$tmp/utf8probe" && "$tmp/utf8probe"

Repository: RustPython/RustPython

Length of output: 218


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "Rust version/options:"
rustc --version
rustc --help 2>/dev/null | grep -E -i "warn|quiet" || true
echo "compile config sample:"
rustc --explain E0658 2>/dev/null | head -5 || true

echo "Utf8Error behavior probe:"
tmp="$(mktemp -d)"
cat > "$tmp/utf8probe.rs" <<'RS'
use core::str::Utf8Error;

fn main() {
    let cases: &[&[u8]] = &[
        b"abc\xff",
        b"abc\xc0",
        b"abc\xc4",
        &[0],
        b"abc\xed\xf3\x99\xbc",
        b"abc\xed\x9f\x80",
    ];
    for bytes in cases {
        let res: Result<&str, Utf8Error> = core::str::from_utf8(bytes);
        if let Err(e) = res {
            let start = e.valid_up_to();
            let hard_end = start + 1;
            let actual_len = e.error_len().unwrap_or(bytes.len() - start);
            let actual_end = start + actual_len;
            println!("bytes={:?} valid_up_to={} hard_end={} actual_len={} actual_end={}", bytes, start, hard_end, actual_len, actual_end);
        } else {
            println!("bytes={:?} ok", bytes);
        }
    }
}
RS
$RUSTC --edition 2021 "$tmp/utf8probe.rs" -o "$tmp/utf8probe" && "$tmp/utf8probe"

Repository: RustPython/RustPython

Length of output: 481


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "Utf8Error behavior probe:"
tmp="$(mktemp -d)"
cat > "$tmp/utf8probe.rs" <<'RS'
use core::str::Utf8Error;

fn main() {
    let cases: &[&[u8]] = &[
        b"abc\xff",
        b"abc\xc0",
        b"abc\xc4",
        &[0],
        b"abc\xed\xf3\x99\xbc",
        b"abc\xed\x9f\x80",
    ];
    for bytes in cases {
        let res: Result<&str, Utf8Error> = core::str::from_utf8(bytes);
        if let Err(e) = res {
            let start = e.valid_up_to();
            let hard_end = start + 1;
            let actual_len = e.error_len().unwrap_or(bytes.len() - start);
            let actual_end = start + actual_len;
            println!("bytes={:?} valid_up_to={} hard_end={} actual_len={} actual_end={}", bytes, start, hard_end, actual_len, actual_end);
        } else {
            println!("bytes={:?} ok", bytes);
        }
    }
}
RS
rustc --edition 2021 "$tmp/utf8probe.rs" -o "$tmp/utf8probe" -Aw lints && "$tmp/utf8probe"

Repository: RustPython/RustPython

Length of output: 300


🌐 Web query:

Rust std::str::Utf8Error error_len valid_up_to documentation

💡 Result:

In Rust, the std::str::Utf8Error struct is returned when an attempt to interpret a sequence of bytes as a UTF-8 string fails [1][2]. It provides two primary methods to help inspect and handle these errors, particularly for incremental decoding or lossy conversions [1][3]. valid_up_to The valid_up_to(&self) -> usize method returns the index in the input byte slice up to which valid UTF-8 was verified [1][3]. It represents the maximum index such that std::str::from_utf8(&input[..index]) would return Ok(_) [1][2]. error_len The error_len(&self) -> Option<usize> method provides additional information about the nature of the failure [1][3]: * Some(len): An unexpected byte was encountered [1]. The value len represents the length of the invalid byte sequence that starts at the index returned by valid_up_to() [1][3]. In scenarios like lossy decoding, you can skip this sequence and resume processing after it (typically after inserting a U+FFFD replacement character) [1][2]. * None: The end of the input was reached unexpectedly [1][3]. This occurs when valid_up_to() is 1 to 3 bytes from the end of the input [1][2]. This often indicates that the input might be a valid UTF-8 character that was truncated (e.g., if the data is being streamed), meaning the sequence could potentially be completed by subsequent chunks of data [1][3]. Both methods are marked as const and are stable [2][3]. These tools are commonly used to implement logic similar to String::from_utf8_lossy without requiring heap allocation [1][4].

Citations:


Use Utf8Error::error_len() for decode-error ranges.

Both host and port from_utf8 errors hard-code the span as valid_up_to + 1, which under-reports invalid multi-byte UTF-8 sequences when building the UnicodeDecodeError; use error_len() or the remaining input length for the end position.

Suggested range calculation
+let start = e.valid_up_to();
+let end = start + e.error_len().unwrap_or(bytes.len() - start);
+
 vm.new_unicode_decode_error_real(
     vm.ctx.new_str("utf-8"),
     vm.ctx.new_bytes(bytes.to_vec()),
-    e.valid_up_to(),
-    e.valid_up_to() + 1,
+    start,
+    end,
     reason,
 )
📍 Affects 1 file
  • crates/stdlib/src/socket.rs#L2615-L2622 (this comment)
  • crates/stdlib/src/socket.rs#L2656-L2664
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/stdlib/src/socket.rs` around lines 2615 - 2622, Update both host and
port UTF-8 error handlers around the host and port from_utf8 calls in
crates/stdlib/src/socket.rs:2615-2622 and crates/stdlib/src/socket.rs:2656-2664
to calculate the UnicodeDecodeError end position using Utf8Error::error_len(),
falling back to the remaining input length when needed, instead of always using
valid_up_to() + 1.

Source: MCP tools

Comment on lines +128 to +136
rustpython_host_env::os::bytes_as_os_str(b).map_err(|e| {
vm.new_unicode_decode_error_real(
vm.ctx.new_str("utf-8"),
vm.ctx.new_bytes(b.to_vec()),
e.valid_up_to(),
e.valid_up_to() + 1,
vm.ctx.new_str("can't decode path for utf-8"),
)
})

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

files=(
  crates/vm/src/function/fspath.rs
  crates/vm/src/stdlib/os.rs
  crates/vm/src/stdlib/posix.rs
)

if rg -n 'valid_up_to\(\)\s*\+\s*1' "${files[@]}"; then
  echo "Found hard-coded one-byte Unicode error ranges" >&2
  exit 1
fi

rg -n 'error_len\(\)' "${files[@]}"

Repository: RustPython/RustPython

Length of output: 416


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '--- fspath excerpt ---\n'
sed -n '110,145p' crates/vm/src/function/fspath.rs | cat -n

printf '\n--- os excerpt ---\n'
sed -n '115,150p' crates/vm/src/stdlib/os.rs | cat -n

printf '\n--- posix excerpt ---\n'
sed -n '1720,1750p' crates/vm/src/stdlib/posix.rs | cat -n

printf '\n--- Utf8Error usages in affected files ---\n'
rg -n 'Utf8Error|bytes_as_os_str|to_str|new_unicode_decode_error_real|valid_up_to|error_len' \
  crates/vm/src/function/fspath.rs crates/vm/src/stdlib/os.rs crates/vm/src/stdlib/posix.rs

Repository: RustPython/RustPython

Length of output: 250


🌐 Web query:

Rust str Utf8Error valid_up_to error_len documentation

💡 Result:

In Rust, the Utf8Error struct is returned when attempting to convert a byte slice into a string slice if the input is not valid UTF-8 [1][2]. It provides two primary methods, valid_up_to and error_len, to help diagnose and handle the error [1][3]. valid_up_to The valid_up_to method returns the index (as a usize) up to which the input string was verified as valid UTF-8 [1][3]. It represents the maximum index such that from_utf8(&input[..index]) would return Ok(_) [1][4]. error_len The error_len method returns an Option that provides more information about the specific failure [1][3]. - Some(len): Indicates that an unexpected byte was encountered [1][2]. The returned length represents the invalid byte sequence starting at the index provided by valid_up_to [1][3]. If performing lossy decoding, you should resume decoding after this sequence, typically after inserting a U+FFFD (REPLACEMENT CHARACTER) [1][3]. - None: Indicates that the end of the input was reached unexpectedly [1][2]. In this case, valid_up_to points to a position 1 to 3 bytes from the end of the input [1][3]. This often occurs when a multi-byte UTF-8 sequence is incomplete because it was cut off (e.g., when reading a stream incrementally) [1][4]. These methods are commonly used together in lossy UTF-8 conversion, as demonstrated in the standard library's implementation of from_utf8_lossy [1][5]. [1][2][3][4]

Citations:


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- fspath excerpt ---'
sed -n '110,145p' crates/vm/src/function/fspath.rs | cat -n

printf '%s\n' ''
printf '%s\n' '--- os excerpt ---'
sed -n '115,150p' crates/vm/src/stdlib/os.rs | cat -n

printf '%s\n' ''
printf '%s\n' '--- posix excerpt ---'
sed -n '1720,1750p' crates/vm/src/stdlib/posix.rs | cat -n

Repository: RustPython/RustPython

Length of output: 4632


Use the complete UTF-8 failure span in all three mappings.

valid_up_to() + 1 only marks one invalid byte. For unexpected EOF, use the source slice length; for invalid overruns, use e.error_len() and handle the None case with the source length.

  • crates/vm/src/function/fspath.rs#128-136
  • crates/vm/src/stdlib/os.rs#133-141
  • crates/vm/src/stdlib/posix.rs#1734-1742
📍 Affects 3 files
  • crates/vm/src/function/fspath.rs#L128-L136 (this comment)
  • crates/vm/src/stdlib/os.rs#L133-L141
  • crates/vm/src/stdlib/posix.rs#L1734-L1742
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/vm/src/function/fspath.rs` around lines 128 - 136, Update the UTF-8
error mappings in crates/vm/src/function/fspath.rs:128-136,
crates/vm/src/stdlib/os.rs:133-141, and crates/vm/src/stdlib/posix.rs:1734-1742
to use the complete failure span: derive the end offset from e.error_len(),
falling back to the source slice length when it is None, and use the source
length for unexpected EOF instead of valid_up_to() + 1.

@moreal moreal added the z-ca-2026 Tag to track Contribution Academy 2026 label Jul 24, 2026

@moreal moreal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your contribution! Could you see the Coderabbit's reviews about Utf8Error::error_len()?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

z-ca-2026 Tag to track Contribution Academy 2026

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants