points by knuckleheads 3 days ago

Funny to see this now. I’ve got a private branch I am in the process of shaping up this weekend to show the compiler team. For deep nested projects like rust analyzer, if you emit the meta data about function types earlier for downstream slots to use, before successful type checking, you can start other crates earlier and use all the slots you have instead of sitting around waiting for the full type checking of the bodies (which other crates largely don’t care about). Something like 40% wall time speed up, maybe 10% to 15% if you have the parallel frontend on.

panstromek 3 days ago

Nice, excited to see that. This is basically deeper pipelining, IIUC. Nick (the author of the blog) implemented the current version of this "emit metadata sooner," optimization and already experimented with pushing it even sooner a while back.

We have discussed this briefly on last All Hands, the problem will probably be dealing with compiler errors, when they happen in builds that already emitted metadata. I believe this was the reason why it wasn't merged in the first place.

  • knuckleheads 3 days ago

    Yes I think I read a post of his from like 2019 were something was tried like this before. My guess is that this will be an interesting and hopefully informative PR for you all to see, but will likely not be merged as is. I had to add the ability for rustc to pause after early meta data emission and that was hacked together I feel, not well designed per se.

    I view this as speculative execution/compilation and just throw out any errors from a crate that depended on another crate that ultimately errors out. It has the same outcome (same errors are printed in both cases) you just have a chance of having totally wasted some cpu time (that was otherwise just sitting around though).

embedding-shape 3 days ago

Sounds lovely, give it a try on the codex-rs which seems to be getting close to using 2K crates in their repository any time now. Takes like 30 minutes to compile on a my beefy machine and it's a goddamn TUI, not sure what's going on, and not too interested in diving into that beast either.

  • knuckleheads 3 days ago

    This thing is a monster, blows out the disk space on a Claude code web session. Will run it on my test server later and let you know.

    • embedding-shape 3 days ago

      Now I want to test it too! Need any testers with Threadripper systems to really see what it can do? ;) Email in profile

      • knuckleheads 3 days ago

        I will email you when I have something up for sure! First cargo check has it at about 28% reduced on codex rs, from 383 seconds down to 277 seconds on a 16 core machine. No -Zthreads yet though, and that typically has a similar effect and reduces the speed up the early meta data gives.

        • knuckleheads 2 days ago

          Result:

          Headstart on codex-rs (16 cores, clean builds, median of 3 alternating off/on pairs):

                                    off       on     saved
            cargo check           383.8s   277.1s    28%
            cargo build           530.7s   330.8s    38%
            check, -Zthreads=8    220.0s   190.5s    13%
            build, -Zthreads=8    317.3s   249.2s    21%
          

          Not helpful for compilation yet, but it's something :)

          • hensenjuang 2 days ago

            Do you see much benefit on incremental builds too, or is this mostly a clean-build win?

            • knuckleheads 2 days ago

              Depends on what you are editing I would say, if it's something deeper in the stack then it should help. The last numbers I have is that this is a 6-8% improvement for incremental editing of the code for cargo itself, but I'm having the benchmark be rerun now, as I've landed a fair amount of changes since then.

  • yearolinuxdsktp 3 days ago

    It’s insane. 45 minutes to build debug from scratch. Linking is super slow even in debug and that’s without LTO.

    Incremental compilation is not cleaned up. Older deps pile up and don’t get cleaned.

    Run out of disk space? Oh it’s just the 200+ GB codex-rs target folder.

    Forget about doing worktrees.

    Rust has a lot of work to do.

    • ninkendo 2 days ago

      > Incremental compilation is not cleaned up. Older deps pile up and don’t get cleaned.

      This is one of those problems that’s both huge/important, and probably impossible to solve. It bites me all the time, I have a measly 1TB nvme disk and I basically have to wipe my target directory and docker build cache (I have a shared cargo cache volume across builds) every other day. I’ve had this workstation for 2.5 years now and I’m honestly worried about nand endurance coming to bite me soon.

      How would one even solve this? When do you know a cache entry won’t be used any more? Trace back provenance of artifacts to their sources, and if the mtimes have been bumped just delete the artifacts? It’s not like resetting source files to older mtimes is a good idea anyway…

      • yearolinuxdsktp 7 hours ago

        I disagree it’s impossible to solve. Huge C++ projects don’t have continuously growing build directories and can still compile incrementally correctly.

        • ninkendo 5 hours ago

          That’s because they overwrite the old artifacts with new ones every time, which has its own problems: you could end up reusing stale artifacts if your compiler version changes, or if different compiler args are used for the same artifact, and a bunch of other stuff. Rust computes a digest of all the relevant build environment settings that may need a different artifact to be built, and uses that digest in the file name. That means you basically never have issues like “oh crap the incremental build broke, do a clean build instead”. (Amongst a whole lot of other benefits I’m sure.)

          But this technique errs on the side of not using invalid artifacts, at the cost of not having a good system for expiring them. To maintain the goal of never having to worry about when to require a clean rebuild, solving cleaning of stale artifacts becomes a much harder problem.

    • WalterBright 2 days ago

      A full debug build of the DMD D compiler takes 3 seconds on my Mac Mini.

    • nicoburns 2 days ago

      > Linking is super slow even in debug

      Linking is normally slower in debug btw (because of the debug info)

      • khuey 2 days ago

        Use the split-debuginfo setting in your Cargo.toml to have the DWARF bypass the linker (on Linux, YMMV on other platforms).

    • kibwen 2 days ago

      If there's a project that takes 45 minutes to produce a debug build from scratch, that's not a Rust problem, that's a them problem.

      Let's compare. I just checked out ripgrep and built it and all its dependencies from scratch. Total time to produce a debug build: 6.92 seconds. This is inside a 16GB VM on a random Linux laptop (on battery); no beefy hardware here.

      Let's try something beefier. I just checked out uutils (an implementation of coreutils) and built it from scratch. Total time to produce debug builds, for eighty binaries and all their dependencies: 30.26 seconds.

      Let's try something beefier. I just checked out Servo, an entire browser engine, and built it and all its dependencies from scratch. Total time to build and link its one-thousand one-hundred and ten compilation units (along with its infamously huge and serial servo-script bottleneck) into a binary: 4 minutes and 38 seconds.

      If Codex takes 45 minutes to build from scratch, that's an indictment of OpenAI and their development processes.

      • surajrmal 2 days ago

        Arguably it is a rust problem. Rust can be improved to reduce the need to do as much manual maintenance of crates dynamics to allow for faster build times. The "you're holding it wrong" answer is not a good perspective to apply and I am surprised to see a prominent member of the rust community using it.

        • kibwen 2 days ago

          No, one can write pathological code in any language. I welcome efforts to improve the performance of the Rust compiler, but nobody is served by optimizing for degenerate use cases to the detriment of legitimate use cases. It is, in fact, possible to hold it wrong; this isn't blithe dismissiveness, but rather an accurate statement of reality, and pretending otherwise doesn't lead to a better language.

    • je42 2 days ago

      This is why I like Zig and Go. :)

    • surajrmal 2 days ago

      Worktrees can work with a shared build cache.

      • embedding-shape 2 days ago

        And this specifically work with Rust and target/ directory? Please do share what you've validated to actually work here.

    • scrubs 2 days ago

      I've abandoned zig for rust. I expected to be annoyed by rust build times, but so far it's been quite good for my needs. To the contrary it has not been any impediment even when prost etc all drag in many deps from a cold build.

      The zig compiler apart from a couple of cross build or linker issues is darn good. Switch to rust is a function of library support, private struct fields. Rust itself is better organized and far better documented compared to zig.

      However, rust has got to get their allocators for container support straightened out soon. I do SAAS for fintech and low level packet transports. Both benefit greatly with allocators custom designed for job at hand.

      45 mins strikes me as something else is amiss ...

      • NobodyNada 1 day ago

        > However, rust has got to get their allocators for container support straightened out soon

        Good news: https://github.com/rust-lang/rust/pull/156882

        The allocator API has been stabilized, to be released in Rust 1.100 in November. (Though only Vec and Box will support custom allocators immediately; the rest of the standard library collections are coming later.)

  • lr1970 2 days ago

    I build it with `cargo build -p codex-cli --bin codex` in under 4 minutes on my old M1 MacBook Pro with 64GB memory compiling more than 1250 crates.

    • embedding-shape 2 days ago

      Sure, that compiles in a few minutes for me too, give it a try with a release build and compiling the full workspace though, since the point here is to stress the compiler :)

  • surajrmal 2 days ago

    It has 1.8 million lines of rust. That compilation time seems close to what I'd expect at that size. If you want something faster you should focus on distributed build caching rather than speeding up the compiler. Also note that crates are effectively equivalent to a single file in c++ as it's the compilation unit. This is similar to compiling 2k files that average 1k loc each in c++. I'm sure it's much worse when you account for macro expansion and generics.

    • mort96 2 days ago

      As they say, it's a damn TUI for a chat bot. 1.8 MLoC is like 50x what you'd expect from something like that. If you want to speed up compilation, reducing that absolutely insane bloat would be a good start

    • embedding-shape 2 days ago

      > If you want something faster you should focus on distributed build caching rather than speeding up the compiler

      ... But the whole point is that that project is big, unwieldy and has lots of crates, both their own and 3rd party. Slapping a cache in front of it kind of defeats the purpose here, to see how parent's thing would speed things up for really large projects.