Hacker Newsnew | past | comments | ask | show | jobs | submit | walkerbrown's commentslogin

I’m curious, what is the objection to Han unification?


Only know the discussion superficially, but from what I understood, the problem was as follows: Chinese Hanzi, Japanese Kanji and Korean Hanja scripts have the same historical roots (Han) and still look very similar when you write them — however, they are fully separate scripts with different use of letters and optical differences that are important.

However, when encoding them into Unicode, the consortium suddenly discovered the importance of being frugal with character IDs and only defined a single "Han" sequence of characters for all three scripts. You need to select a matching font to actually display the text in the script that's intended.

This makes using those characters an unnecessary hassle and also came over as a cultural faux-pas to some (essentially like encoding Greek Eta, Latin H and Cyrillic En as the same character, because they look similar and the scripts have a common history)

<s>It also looks a bit hypocritical when Unicode otherwise goes the opposite way and has the space to encode dozens of variants of the Latin alphabet or — as we see here — several categories of dashes.</s> (Edit: Han unification was done when Unicode was still 16 bit, so there were real space constraints — so no hypocrisy actually)

See here: https://en.wikipedia.org/wiki/Han_unification


I honestly agree with the arguments they made to defend it, but then they tied themselves in loops trying to defend things like separate Greek and Cyrillic scripts and making superfine distinctions about how that isn't the same phenomenon. But they couldn't merge those scripts together now even if they wanted to, because of the stability policy. And the existence of the blackletter, small-caps, superscript, etc. characters seems entirely indefensible to me. All font rendering issues; it's clearly the same symbol with the same fundamental meaning.

Really, they needed to come up with the idea of variation selectors way earlier, and use them aggressively. And design something more consistent for character composition, and skip the precomposed characters (those really are glyphs representing a cluster; being able to treat the cluster as an atomic entity versus its components doesn't add actual semantic information to the text).


If I'm writing a text about Russian language in English, I can easily include an example in Russian: "привет, как дела?" and it will be displayed correctly.

But, if I was writing about Mandarin Chinese in Japanese, I just wouldn't be able to do it using a simple text input (like HN comment form). I would need more advanced formatting tools to make sure that both Chinese and Japanese characters are displayed correctly.

Because, as far a I understand, one Unicode symbol could correspond to both a Hanzi character and a Kanji character. Which character will be displayed is only determined by font. Since HN doesn't allow me to select different fonts for different parts of my comment, it is simply impossible to create a comment that includes both Japanese and Chinese characters.


Uber misclassifies en masse the drivers it employs and shifts vehicle depreciation costs onto these "small businesses". Our taxes pay for healthcare costs which should rightfully be borne by the drivers' employer. I would say that's a non-negligible social cost.


It's weird to describe this as "employment" when many and perhaps most drivers these days are simultaneously "employed" by both Uber and Lyft, serving rides for both. This kind of informal relationship is a lot closer to contracting than traditional employment.


Is it quacking like employment? Sort of and sort of not.

Not: very flexible hours, not exclusive, driver can switch customer? driver provides tools

Is: driver cannot subcontract, has to follow set rules (maybe that makes is franchise like?)

Of course they took something that was employment and made it not employment at scale.


It’s perfectly normal to have two W-2 employers.

On the other hand, it’s odd that these “contractors” can’t set prices for their services and risk deactivation if they decline the price Uber offers.


Some very recently published research (Dec 2025) claims evidence of fire starting among homo neanderthalensis. This would push back fire starting know-how (not only control) from 50k to 400k years ago. Cool stuff!

[1] https://www.nhm.ac.uk/press-office/press-releases/groundbrea...


For anyone this deep on the thread, check out this video (great presenter!) explaining TV spectrum allocation, NTSC, PAL, and the origin of 29.97 fps.

https://youtu.be/3GJUM6pCpew


TIL NTSC: He explained that NTSC stands for Not The Smartest Choice, but I always assumed it meant Never The Same Color.


Nice work OP, and congrats on HN front-page. Keep publishing or it never happened!


Improved transcription accuracy from non-native English I bet.


No snark, but posting with genuine concern for any homebuyers who didn’t experience the pain of 2008, please read and grok this: https://www.investopedia.com/terms/u/underwater.asp


I’m in favor of cycling infrastructure in the US, but it’s important to remember that the latitude of the Netherlands (like most of the European cycling cities) is north of the entire continental US.

I’m in decent shape, but will sweat at a resting heart rate in a Southern US summer.


Yes, I was definitely expecting cuneiform and base 60 literals.


I want this now. Bonus points if it's a vector language and is as terse as K.


I thought the failure was interesting too, enough to try on GPT-4. It succeeds with the same prompt.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: