Hacker Newsnew | past | comments | ask | show | jobs | submit | someone_19's commentslogin

> because of essential constraints of low-level languages that prevent them from doing certain optimisations that matter mostly in large programs

Which specific optimizations are you referring to?

In my experience, this is largely a myth; compared to Rust, you actually get even faster code right away.

JIT is effective for languages where the source code lacks sufficient information (dynamic typing, where anything can be null).


> Which specific optimizations are you referring to?

A JIT with speculative optimisation and a moving GC.

There are two constraints in low-level languages that trump any of their performance goals, one technical and one a matter of preference.

The technical limitation is that they must use stable pointers (because they need to be low-level and so having an FFI layer that separates "hardware pointers" from "language pointers", as we have in Java defeats their main purpose). This means that you need to translate data storage or code storage to hardware addresses, and that interferes with both moving collection and with JIT compilation.

The other constraint is that low-level languages value worst-case performance over the average-case and even amortised performance. These languages prefer an operation (e.g. dynamic dispatch) to be slow as long as it's never too slow. With a JIT (and I describe more later), virtual dispatch can be super-fast almost all the time, but occassionally, you'll hit a trap because the speculation was wrong, and then you need to deoptimise and recompile.

> In my experience, this is largely a myth; compared to Rust, you actually get even faster code right away.

We wouldn't be doing it in the first place if it was a myth. In a low-level language, you can get very fast code if you do some manual optimisations, but they don't easily scale as the program grows and evolves, because they're viral. The two most basic examples are dynamic dispatch (which is the most general mechanism, which scales the best in terms of program evolution) and shared heap objects (again, the most general mechanism). These become more common and less easily avoided over time, and they're slow in low-level languages because of the constraints I mentioned.

That low-level languages make it harder and harder to preserve good performance over time as they evolve and grow is a problem familiar to those who've worked for years on large software written in a low level language (as I have). The JVM was designed, among other things, to solve this performance problem in large programs.

> JIT is effective for languages where the source code lacks sufficient information (dynamic typing, where anything can be null).

A JIT can make such languages decently fast, but that's not how it's used in Java. In Java it is used for speculative optimisation, which allows far more aggressive optimisation than an AOT compiler can do. E.g. by default, Java inlines and specialises virtual calls 15 levels deep. An AOT compiler can't do that or its code will explode. We get around it with selective use of templates in C++ (or comptime in Zig), but it has to be selective, and it's viral.


Thank you for the reply.

Do you mind a reasoned discussion?

> A JIT with speculative optimisation and a moving GC.

Idiomatic Rust, through its concepts of ownership and borrowing, encourages a pattern where you receive data as an argument or create it directly, perform operations on it, and then discard it via RAII. This bears some resemblance to functional programming. This approach does not apply to buffers of unknown size, which still require heap allocation; unfortunately, Rust lacks automatic buffer reuse. However, such optimization is theoretically possible. The stack is definitely faster than anything else.

> This means that you need to translate data storage or code storage to hardware addresses, and that interferes with both moving collection and with JIT compilation.

You don't need GC if you allocate data on stack. You also do not need to dereference the pointer.

> dynamic dispatch

You mentioned templates. In Rust, traits that are monomorphized - much like templates-are the standard approach; using vtables or `dyn trait` is a relatively rare use case. This stems from the fact that all code is known at compile time and there is no dynamic loading, allowing the compiler to eliminate polymorphism from the code entirely.

> and shared heap objects

This might be considered convenient, but in my view, it also leads to code that is harder to maintain when objects can be modified from multiple places. However, I think that is outside the scope of the current discussion.

> We get around it with selective use of templates in C++ (or comptime in Zig), but it has to be selective, and it's viral.

Yes, monomorphization is the default solution in Rust. It is not always viral either, because when using it, you often define specific types, and they do not spread beyond that scope.

I suppose you could say that the programming style I am talking about is complex, inconvenient, unmaintainable, and so on. What I mean is, assuming this programming style is sufficiently convenient—and perhaps even has its own advantages - then none of the optimizations you listed offer an edge, and the Rust code will definitely be faster.


> The stack is definitely faster than anything else

I have seen it mentioned everywhere, but is this actually true?

I mean, of course it is faster than random cold memory, but is it actually faster than a hot, in-cache part of the heap? It is not special in any other way, AFAIK.

And for what it's worth, what pron mentioned, Java uses a pretty similar structure for initial allocation, a thread local buffer where you just pointer bump. Another thread can then in the background copy still alive objects from this "arena" and then reset the whole thing.


> I have seen it mentioned everywhere, but is this actually true?

Yes, it just adding or subtraction int to stack pointer register. I’m not certain, but the only thing that might be faster is accessing data at a fixed address - that is, global variables.


That's the way of getting the address itself, that's unrelated to how fast the actual memory read/write is.

Stack is fast because it is frequently "touched" staying in cache. If you were to continuously read write a small segment of the heap, I don't think it would fair any worse than "the stack". This was my point


> However, such optimization is theoretically possible. The stack is definitely faster than anything else.

What you're describing isn't a stack, but an automatic arena, and this optimisation is easier to do in Java. It's easier to do in Java because it requires setting a "current arena" or inlining, both of which Java can do more easily, and then either the arena will be heap allocated (which will be slower in Rust) or associated with the thread, which is not something low-level languages tend to do.

> You don't need GC if you allocate data on stack. You also do not need to dereference the pointer.

Moving collectors don't need to dereference anything (they don't know and don't want to know when an object is "dead"), and stack allocation works in both languages, only, as you pointed out, is not quite general (not every data structure with a known lifetime can be allocated on the stack).

> You mentioned templates. In Rust, traits that are monomorphized - much like templates-are the standard approach; using vtables or `dyn trait` is a relatively rare use case. This stems from the fact that all code is known at compile time and there is no dynamic loading, allowing the compiler to eliminate polymorphism from the code entirely.

Sure, except Java does this automatically, and it can do it more aggressively. Dynamic dispatch is rare in low-level languages because it's expensive in those languages. But it's not easy to avoid as programs get larger. That is exactly one of the problems in large programs that the JVM set out to solve.

> This might be considered convenient, but in my view, it also leads to code that is harder to maintain when objects can be modified from multiple places. However, I think that is outside the scope of the current discussion.

I agree that whether it has downsides is outside the scope of this discussion, but the point is that as programs evolve and grow, the abstractions tend to be more general, and low-level languages suffer from "abstraction cost", where the more general abstraction (which becomes more common over time) is more expensive. Again, this is exactly why large C++ programs suffered from performance issues and what the JVM tried to address.

> Yes, monomorphization is the default solution in Rust.

... and in C++. But it is viral, and Java monomorphises without suffering from "zero overhead abstractions".

The ability to move pointers, both to data and to code, opens up the possibility of using JITs and moving GCs, which are very powerful optimisations. A JIT does impose two further tradeoffs (aside from the need for an FFI layer), though, which are warmup and the possibility of deoptimisation. We can now cache the generated machine code from one execution to another (https://openjdk.org/jeps/544), but the possibility of deoptimisation remains (in fact, it's what enables the aggressive speculative optimisations), which means you gain average (or even amortised) performance at the cost of the worst case.

Anyway, the JVM was designed as a solution for the performance issues low-level languages suffer from as programs grow and/or evolve. It comes with tradeoffs, but those most affect small or short-lived programs.

The thing to remember is that low-level languages are not optimised for performance but for low-level control (i.e. pointers are direct addresses etc.). Such control can translate to good performance when programs are small (see next) but it becomes a practical hindrance to performance when they're large.

> I suppose you could say that the programming style I am talking about is complex, inconvenient, unmaintainable, and so on. What I mean is, assuming this programming style is sufficiently convenient—and perhaps even has its own advantages

That advantage is a performance advantage. The question isn't "does there exist (in the mathematical sense) some program that is fast?" but "how fast is the program we can write within the budget we have?" When programs are small, manual optimisation is practical; when they grow large - not so much. And that's excluding the matter of a moving collector, which is just hard to compete with on speed regardless of program size, unless you use areans, but they're not at all easy to use in most low-level languages except Zig.

> and the Rust code will definitely be faster.

This is true only in the abstract mathematical sense. The reason we don't write programs that we want to be fast in Assembly (which is faster than anything in the same sense: for any program in any language, there exists and Assembly program that's at least as fast) is not because other languages are fast enough, but because in practice the programs we can actually write in the budget we have will be faster than the Assembly programs we could write. Of course, that could change when AI is able to generate perfect low-level code, but when that happens, it might as well generate machine code directly.


> Assembly (which is faster than anything in the same sense: for any program in any language, there exists and Assembly program that's at least as fast)

At least you aren't claiming that the JVM is ~1.5 faster than perfectly written assembly :)

I disagree with a lot of what you’re writing. However, we’ve reached the point where we need to run benchmarks and analyze the generated code (this is easy to do for compiled languages using https://godbolt.org/, but for the JVM, it can be a bit more complex, given the warm-up factor).

So, there is one fundamental point I started with:

> JIT is effective for languages where the source code lacks sufficient information (dynamic typing, where anything can be null)

And your answer is:

> A JIT can make such languages decently fast, but that's not how it's used in Java. In Java it is used for speculative optimisation, which allows far more aggressive optimisation than an AOT compiler can do.

Essentially, you are saying that the compiler can apply aggressive optimizations when it knows what is happening in the code.

But I say that JIT is needed so the compiler can figure out what is happening in the code and perform aggressive optimizations.

There are many things that can be inferred from the code without needing to execute it. The question is how difficult it is to make such an inference: in one scenario, the compiler might attempt to track whether specific data changes-and, if it can prove this, mark the data as immutable and apply certain optimizations-whereas in another, it might already possess the information that the data is immutable.

Moreover, information about immutability is useful not only to the compiler but also to the programmer. Just like information about types: it benefits both the compiler and the programmer. Imagine a fan of JS or Python joining our conversation and claiming that both Java and Rust are low-level languages because you have to specify types - something they view as complex and a hindrance to development speed.

The same applies to the GC: the compiler can perform more optimizations when it knows when memory needs to be cleared (move it to stack or even place the data on registers). The JVM attempts to do this (via escape analysis), but there are limitations; consequently, data ends up on the heap, and GC operations come at a cost (due to data movement).

Rust simply makes it easy to obtain far more information, enabling aggressive optimizations that are both immediate and guaranteed.

There remain a small number of cases, such as `switch` statements - where one branch executes 99% of the time, while the other 99 branches execute only 1% of the time. In such instances, the JIT could indeed perform further optimizations; however, I am not even sure if the overhead of monitoring wouldn't outweigh the benefits. And the question is when and how to perform PGO, or whether to perform it at all.


> There are many things that can be inferred from the code without needing to execute it. The question is how difficult it is to make such an inference: in one scenario, the compiler might attempt to track whether specific data changes-and, if it can prove this, mark the data as immutable and apply certain optimizations-whereas in another, it might already possess the information that the data is immutable.

Yes, and the important point is that when it comes to knowing things statically, abstraction and optimisation are in conflict. The whole point of abstraction is that the implementation details aren't known. So in C++ we always suffer from this problem called "zero overhead abstractions" or "abstraction costs", which means that to give the compiler the information it needs, we have to use less general abstractions, which are viral and harm evolution. What a JIT does is allow the compiler to learn the very things that abstraction hides; yes, it's a virtual call, yes, it could target anything, but I've seen it hit the same target 1000 out of the last 1000 times, so I speculate that this will continue and I'll inline even though I could be wrong.

> The same applies to the GC: the compiler can perform more optimizations when it knows when memory needs to be cleared

I understand why this could be true in theory, but in practice the problem is:

1. not that the compiler knows when an object is unreachable, but that the generated code has to do something at that point, and

2. the most efficient known memory management algorithms - moving collectors and arenas, both work in nearly the same way - are entirely predicated on freeing memory in bulk and on not doing anything when an object becomes unreachable, and so the knowledge of when an object becomes unreachable doesn't help them.

So it is true that C and C++ and Rust always statically know when an object is dead, and you could say that hypothetically they don't need to do anything with that information, but in practice they all act on that information immediately and that's inefficient.

> There remain a small number of cases, such as `switch` statements - where one branch executes 99% of the time, while the other 99 branches execute only 1% of the time.

So the main practical benefit of a JIT isn't that at all, but that it can do the "mother of all optimisations" - inlining - far more aggressively. Inlining is important because it cracks open the abstraction boundary of the inlined subroutine, and allows the compiler to further specialise and optimise things, now with the appropriate context.

Anyway, all of these fundamental questions and differences between languages with more statically known information and figuring out "unprovable" information in practice were very well known before the JVM was built to address the performance problems we had suffered from in large C++ programs. So we can argue over which workloads are helped by this and which aren't, but there is no way to say which is usually faster in the absract (because, again, these considerations were known and taken into account). It's merely an empirical question, and not one that's easy to settle. After more than 25 years of working with C++ and almost 20 years of working with Java, my default is that low-level wins on performance (if written by experts) in smaller programs, and Java wins on performance in larger programs, but of course, there are many caveats in either direction.


> The whole point of abstraction is that the implementation details aren't known.

I disagree with that phrasing; it is better to say that abstractions allow a programmer to ignore unimportant details. For example, when developing two modules (possibly even by different teams), all they know about each other is a lean interface, without any implementation details. However, the compiler might know everything.

> So in C++ we always suffer from this problem called "zero overhead abstractions" or "abstraction costs"

This is another odd term. In Rust, the term used instead is "zero-cost abstractions," referring to cases where the compiler can generate instructions for higher-level code just as efficiently.

> So the main practical benefit of a JIT isn't that at all, but that it can do the "mother of all optimisations" - inlining - far more aggressively.

I’ll reiterate that I disagree with this: inlining is performed very efficiently during monomorphization. And monomorphization is used very frequently in Rust.

> After more than 25 years of working with C++

I don't have much experience with C++; I mostly use Rust. I can only assume that the C++ development experience is far worse than Rust - especially when trying to write software that is both reliable and fast. This may be particularly relevant to older C++.

So, a Dog, a Cat, and an Abstract Mammal walk into a bar...

https://godbolt.org/z/dEsW1sfM8

I didn't want to do this, but I went ahead and created a small example showing that monomorphization and inlining work remarkably well. (Obviously, this example does not address memory management)


> For example, when developing two modules (possibly even by different teams), all they know about each other is a lean interface, without any implementation details. However, the compiler might know everything.

You're talking about abstraction at the code level; I'm talking about abstraction at the language level. A virtual call means "the implementation is unknowable here", and it is, indeed, rarely knowable to an AOT compiler.

> This is another odd term. In Rust, the term used instead is "zero-cost abstractions," referring to cases where the compiler can generate instructions for higher-level code just as efficiently.

Rust took that term from C++ (and it had slightly different ones over the years). What it means that the language offers different mechanisms - chosen statically - with different abstraction levels (i.e. different generality) and different costs, some of which are zero, but often similar or identical-looking code at the use site, because the mechanism choice depends on some non-local information, typically associated with the type. I call it "writes like a low-level language, reads like a high-level one". This is different from C (or Zig), which usually makes the selected mechanism explicit at the use site, or from Java, which chooses the cheapest applicable mechanism at every use-site for a single general construct.

The problem is that, because the mechanims is chosen statically, you need to choose the cheapest applicable mechanism, usually virally, yourself, and that over time this gets harder or things drift toward the more general and costly mechanisms.

That's what Java tried to solve, but there are, of course, tradeoffs. The obvious one (which has solutions) is warmup time, because the compiler needs to wait to learn what optimisations can be applied even if they're unprovable, e.g. to learn that a polymorphic application is actually monomorphic in practice at a particular call-site (the solution is to cache the optimised machine code from one run to the next). The more fundamental tradeoffs are 1. you're not guaranteed which mechanism will be chosen, 2. there can be a bad, though amortised, worst-case due to deoptimisation (this is what happens when the compiler optimises too aggressively and then finds out it was wrong, e.g. it inlined a virtual call under the assumption it's the only target at the use site, but after a while, another target appears (in Rust/C++, you'll always pay the higher price, but there's no point at which deoptimisation occurs), and 3. you need an FFI layer, as you can't take the machine address of a compiled subroutine (as it may be re-compiled multiple times).

Tradeoffs 2 and 3 are the main reasons low-level languages don't do this optimisation, and 3 is particularly important. Low-level languages are designed, first and foremost, to be low level. To do its sophisticated optimisations, Java needs to move around pointers to both code and data, which requires a clear FFI layer between Java code and anything external. Having such an FFI layer in a low-level language (and I'm not talking about Rust/C++'s thin extern FFI) defeats the very purpose of a low-level language, which is to talk directly to the hardware and OS. That is the chief goal of all low-level languages, and they sacrifice everything for it. Not only safety (Rust's unsafe is used relatively pervasively) but also performance.

> I can only assume that the C++ development experience is far worse than Rust

Actually, the experience in the two languages is remarkably similar, and not by accident. Rust certainly improves some details, but the overall experience "in the large" is very close. But note that the performance problem is not because of "zero cost abstractions" but because of the low-levelness and focus on the worst-case. Even in Zig, which tries hard to avoid zero cost abstractions to keep use sites explicit, the choice between a specific-and-cheap and a general-and-expensive mechanism means that for best performance you need to pick a less general mechanism, and that gets trickier and trickier to maintain as the program evolves over the years, and especially if it's large.

> monomorphization and inlining work remarkably well.

Of course it does, which is why the optimising JIT was invented: to make it work more broadly!

This wasn't done just on principle, but to solve a very real problem. What we used to do in C++ is architect a solution and write code that monomorphises in all the right places - because that's what one does - and the result was good and fast. And then, five years later, we had to add some feature and were faced with the choice of either undoing some core optimisation or re-architecting some 10,000 LOC. The problems didn't arise when first writing the program, when everything was known. It arose when some change - that hadn't been foreseen when the program was first written - had to be done. Java didn't make the first step substantially cheaper; it made all the following work - five, ten, fifteen years down the line - substantially cheaper.

An important caveat is that HotSpot currently misses many auto-specialisation opportunities that it could take advantage of, but that's one of the things that make working on such a cutting-edge compiler so interesting :) The problem, as always, isn't just the work required, but also determining which optimisations actually make a difference in real programs (and not just in specific benchmarks).

Of course, now there's this hypothesis that AI could do this costly rearchitecting for you, even in large programs, but it doesn't do it well (at all!) today, and I think that when we get to a point where it can do it well, it will also be smart enough to do it in machine code directly (or at least in C), at which point all programming languages will be over. What I don't think is likely is that AI will be able to do extremely complex semantics-preserving large-scale transformations correctly, yet still need the help of a sophisticated compiler for much more local transformations and far simpler correctness checks.


> No, you’re clearly wrong; golang was always going to add support for generic functions.

Let's get this straight. I'll give you a long quote from Rob Pike's article where he describes the history of the go language:

""" One thing that is conspicuously absent is of course a type hierarchy. Allow me to be rude about that for a minute.

Early in the rollout of Go I was told by someone that he could not imagine working in a language without generic types. As I have reported elsewhere, I found that an odd remark.

To be fair he was probably saying in his own way that he really liked what the STL does for him in C++. For the purpose of argument, though, let's take his claim at face value.

What it says is that he finds writing containers like lists of ints and maps of strings an unbearable burden. I find that an odd claim. I spend very little of my programming time struggling with those issues, even in languages without generic types.

But more important, what it says is that types are the way to lift that burden. Types. Not polymorphic functions or language primitives or helpers of other kinds, but types.

That's the detail that sticks with me.

Programmers who come to Go from C++ and Java miss the idea of programming with types, particularly inheritance and subclassing and all that. Perhaps I'm a philistine about types but I've never found that model particularly expressive.

My late friend Alain Fournier once told me that he considered the lowest form of academic work to be taxonomy. And you know what? Type hierarchies are just taxonomy. You need to decide what piece goes in what box, every type's parent, whether A inherits from B or B from A. Is a sortable array an array that sorts or a sorter represented by an array? If you believe that types address all design issues you must make that decision.

I believe that's a preposterous way to think about programming. What matters isn't the ancestor relations between things but what they can do for you.

That, of course, is where interfaces come into Go. But they're part of a bigger picture, the true Go philosophy. """

Rob Pike, 2012

I can draw a few conclusions from this: firstly, he didn't want to add generics at all because he didn't think they were useful, and secondly, he doesn't understand programming very well and doesn't know what generics are and confuses them with inheritance.


You already have iota. Type safety is not needed by design:

> Go intentionally has a weak type system... Go in general encourages programming by writing code rather than programming by writing types...

https://github.com/golang/go/issues/29649#issuecomment-45482...


Theres a whole universe between proper enums and go becoming Ada. Same with null safety.


Ada is the lost knowledge of a past, more advanced civilization.



...and Java didn't even have basic enums or sum types from the beginning. But it had null.

They added enums, they added sealed classes. They're trying to get rid of null (apparently it's really hard). The problem is that in 2012, when go 1.0 was released, this should have been obvious to everyone.

Here's a famous discussion from 2009, three years before the 1.0 release (tldr: facepalm)

https://groups.google.com/g/golang-nuts/c/rvGTZSFU8sY


I remember back in 1995 thinking that it was stupid for Java not to have generics, so instead you had to always cast Vector/Hashtable elements from Object, or implement your own type-specific container classes for every element type (and there wasn’t even a preprocessor to facilitate the latter).

Sum types I didn’t really miss, because you can implement a type-safe equivalent using the Visitor pattern, and retain an interface-implementation separation that native sum types typically don’t provide.


That is an interesting read, seems some were having a hard time grasping the benefit of having compiler checks for potential null dereferences. Having worked with null safety in TypeScript and Kotlin the extra bit of strictness is nice.


The biggest design issue with adding value types and non nullable to Java, is that the number one design requirement of any solution is not to break Maven Central.

Every compiled JAR out there has to keep working as always on a JVM with updated semantics, and worse code has to be compatible, when passing class instances around between old and new code.

Then there are the guest languages on the JVM as well.


So you mean to say that PHP5 and Js from 2007 had a well-founded design?


I agree that they were clearly not in a hurry. I disagree that they are doing everything right. I am interested to see how they will fix the 'million dollars mistake'.


Indeed, in 2012, it was not clear to anyone that generics were needed /s


> but the good parts of JS are better than anything else in existence

What you talking about?! I can't think of a single thing in Js that I could say is good.

Okay, two big corporations have invested a lot of money and effort into making V8 and TypeScript, and now it's useful. But I don't consider it exactly part of Js.


Unhappy way is a part of contract. So yes, that is what I want. If a function couldn't fail before but can after the update - I want to know about it.

In fact, at each layer, if you want to propagate an error, you have to convert it to one specific to that layer.


You can do this quite easily in Rust. But you have to overload operators to make your type make sense. That's also possible, you just need to define what type you get after dividing your type by a regular number and vice versa a regular number by your type. Or what should happen if when adding two of your types the sum is higher than the maximum value. This is quite verbose. Which can be done with generics or macros.


You can do it at runtime quite easily in rust. But the rust compiler doesn’t understand what you’re doing - so it can’t make use of that information for peephole optimisations or to elide array bounds checks when using your custom type. And you don’t get runtime errors instead of compile time errors if you try to assign the wrong value into your type.


Here is example of compile time error with wrong newtype argument:

https://play.rust-lang.org/?version=stable&mode=debug&editio...

rust-analyzer gives an error directly in IDE.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: