Skip to content

Warn before opening a file that looks like binary content as text - #763

Open
KAMI911 wants to merge 1 commit into
linuxmint:masterfrom
KAMI911:build/perf/warn-before-opening-binary-file
Open

KAMI911 wants to merge 1 commit into
linuxmint:masterfrom
KAMI911:build/perf/warn-before-opening-binary-file

Conversation

@KAMI911

@KAMI911 KAMI911 commented Oct 3, 2026

Copy link
Copy Markdown

Motivation

Opening a binary file (an image, an executable, an archive, etc.) as text used to be extremely slow and got worse than linearly with file size: measurements showed a 1 MB file with binary content taking about 21 seconds to load and settle, versus about 2.5 seconds for a normal 1 MB text file, and a 5 MB binary file did not finish loading even after 45 seconds. The cause is that content which cannot be cleanly decoded as text makes the file loader try each candidate character encoding in turn before falling back to a lossy conversion, and the resulting buffer of mostly unreadable characters is also expensive for the text view to lay out.

This adds a quick check before loading starts: the first 8 KB of the file are read and scanned for a NUL byte, the standard heuristic also used by tools like git and diff to tell binary content from text. If a NUL byte is found, a dialog warns that the file does not look like a text file and that opening it as text can be slow and will show unreadable content, letting the user cancel or open it anyway. Loading proceeds exactly as before once confirmed, or for any file where no NUL byte is found in that initial chunk.

Result

Installed the built .deb and drove the real binary headlessly (Xvfb + xdotool/xprop), comparing a crafted binary file against a normal text file:

  • The binary file reliably spawns a second, distinct modal window (629×129, WM_CLASS=Xed) and the tab never finishes loading while it's up — the main window title stays at the default "xed" instead of switching to the file name.
  • The text file shows no such window; the main window title updates to the document name immediately, same as before this change.
  • The detection heuristic itself (looks_like_binary_file) was also unit-tested in isolation against real samples — a C source file, a dictionary word list, an empty file, /bin/ls, 2 MB of random bytes, and a text file whose only NUL byte sits just past the 8 KB sniff window (expected to not be flagged, by design) — all 6 cases detect correctly.

Opening a binary file (an image, an executable, an archive, etc.) as text used to be extremely slow and got worse than linearly with file size: measurements showed a 1 MB file with binary content taking about 21 seconds to load and settle versus about 2.5 seconds for a normal 1 MB text file, and a 5 MB binary file did not finish loading even after 45 seconds. The cause is that content which cannot be cleanly decoded as text makes the file loader try each candidate character encoding in turn before falling back to a lossy conversion, and the resulting buffer of mostly unreadable characters is also expensive for the text view to lay out. This adds a quick check before loading starts: the first 8 KB of the file are read and scanned for a NUL byte, the standard heuristic also used by tools like git and diff to tell binary content from text. If a NUL byte is found, a dialog warns that the file does not look like a text file and opening it as text can be slow and will show unreadable content, and lets the user cancel or open it anyway. Loading proceeds exactly as before once confirmed, or for any file where no NUL byte is found in that initial chunk.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant