Nп/п : 26 из 100
От : Eugene Subbotin 2:5075/35 30 авг 26 05:04:32
К : Michiel van der Vlist 30 авг 26 05:13:01
Тема : Re: Name in Cyrrilic
----------------------------------------------------------------------------------
@REPLY: 2:5075/35@fidonet 6a937f20
@MSGID: 2:5075/35@fidonet 6a93911c
@CHRS: UTF-8 4
@TZUTC: 0300
@TID: hpt/nbsd 1.9 2024-03-02
Hello Michiel!
Sunday August 30 2026 03:38, I wrote to you:
ES> Your message suddenly brought to light a huge problem. And the problem
ES> isn╩╝t with the reader, but with the tosser. The thing is, the original
ES> message with a cyrillic UTF-8 name in the To: field, for some reason,
ES> didnтАЩt make it into my JAM database for this echo.
ES> However, it passed through my station and reached 2:5075/21 without
ES> any issues. There╩╝s clearly a problem with the code of the hpt tosser
ES> from the Husky project: I╩╝ve already run into this once beforeтАФit has
ES> issues with cyrillic names in message headers, even in CP866. All in
ES> all, this will require further study of the issue.
ES> This shows that the technical readiness to use non-ASCII names in the
ES> тАЬFromтАЭ and тАЬToтАЭ fields is still limited and can cause problems even
ES> with fairly modern software.
Following up on that point, a few more thoughts came to mind.
Not only did my tosser, for some reason, fail to store a record
with the To-field тАЬ╨Х╨▓╨│╨╡╨╜╨╕╨╣ ╨б╤Г╨▒╨▒╨╛╤В╨╕╨╜тАЭ in the JAM database
тАФ and the reasons for this are still unknown тАФ but there may be
even more problems overall.
For example, the name тАЬ╨Р╨╗╨╡╨║╤Б╨░╨╜╨┤╤А ╨е╤А╨╕╤Б╤В╨╛╤Д╨╛╤А╨╛╨▓тАЭ in
UTF-8 will take up 41 bytes, which wonтАЩt fit within the 36 bytes
defined in FTS-0001. This means that five bytes will be truncated along the
way, which will corrupt the UTF-8 text. With GoldED+, exporting truncation
will occur correctly тАФ by character rather than by byte тАФ but it
may work differently on other systems.
So there are far more problems, and they require a technical
solution and standardization first before we can use such fields from the
UTF-8 nodelist in headers.
For example, FSP-1030 proposed using the kludges ^AUCSFROM:, ^AUCSTO:,
and ^AUCSSUBJ: to include UTF-8 header fields in messages, but this was
not adopted as a standard.
Eugene
... It`s full of stars!
--- GoldED+/BSD 1.1.5-b20260829 (NetBSD 11.0 Intel Core Haswell)
* Origin: FireFox Station (2:5075/35)
SEEN-BY: 154/10 203/0 240/5832 280/464 5003 5555
292/789 301/1 310/31 341/66
SEEN-BY: 460/58 5001/100 5015/46 5019/40 5020/715
1042 1146 5452 9696 5023/24
SEEN-BY: 5030/1081 5051/44 5075/21 35 6035/3
@PATH: 5075/35 280/5555 5020/715