- UTF-8 ------------------- < Пред. | След. > -- < @ > -- < Сообщ. > -- < Эхи > --
 Nп/п : 26 из 100
 От   : Eugene Subbotin                     2:5075/35         30 авг 26 05:04:32
 К    : Michiel van der Vlist                                 30 авг 26 05:13:01
 Тема : Re: Name in Cyrrilic
----------------------------------------------------------------------------------
                                                                                 
@REPLY: 2:5075/35@fidonet 6a937f20
@MSGID: 2:5075/35@fidonet 6a93911c
@CHRS: UTF-8 4
@TZUTC: 0300
@TID: hpt/nbsd 1.9 2024-03-02
Hello Michiel!

Sunday August 30 2026 03:38, I wrote to you:

 ES> Your message suddenly brought to light a huge problem. And the problem
 ES> isn╩╝t with the reader, but with the tosser. The thing is, the original
 ES> message with a cyrillic UTF-8 name in the To: field, for some reason,
 ES> didnтАЩt make it into my JAM database for this echo.

 ES> However, it passed through my station and reached 2:5075/21 without
 ES> any issues. There╩╝s clearly a problem with the code of the hpt tosser
 ES> from the Husky project: I╩╝ve already run into this once beforeтАФit has
 ES> issues with cyrillic names in message headers, even in CP866. All in
 ES> all, this will require further study of the issue.

 ES> This shows that the technical readiness to use non-ASCII names in the
 ES> тАЬFromтАЭ and тАЬToтАЭ fields is still limited and can cause problems even
 ES> with fairly modern software.

Following up on that point, a few more thoughts came to mind.

 Not only did my tosser, for some reason, fail to store a record
with the To-field тАЬ╨Х╨▓╨│╨╡╨╜╨╕╨╣ ╨б╤Г╨▒╨▒╨╛╤В╨╕╨╜тАЭ in the JAM database
тАФ and the reasons for this are still unknown тАФ but there may be
even more problems overall.

 For example, the name тАЬ╨Р╨╗╨╡╨║╤Б╨░╨╜╨┤╤А ╨е╤А╨╕╤Б╤В╨╛╤Д╨╛╤А╨╛╨▓тАЭ in
UTF-8 will take up 41 bytes, which wonтАЩt fit within the 36 bytes
defined in FTS-0001. This means that five bytes will be truncated along the
way, which will corrupt the UTF-8 text. With GoldED+, exporting truncation
will occur correctly тАФ by character rather than by byte тАФ but it
may work differently on other systems.

 So there are far more problems, and they require a technical
solution and standardization first before we can use such fields from the
UTF-8 nodelist in headers.

 For example, FSP-1030 proposed using the kludges ^AUCSFROM:, ^AUCSTO:,
and ^AUCSSUBJ: to include UTF-8 header fields in messages, but this was
not adopted as a standard.

Eugene

... It`s full of stars!
--- GoldED+/BSD 1.1.5-b20260829 (NetBSD 11.0 Intel Core Haswell)
 * Origin: FireFox Station (2:5075/35)
SEEN-BY: 154/10 203/0 240/5832 280/464 5003 5555
292/789 301/1 310/31 341/66
SEEN-BY: 460/58 5001/100 5015/46 5019/40 5020/715
1042 1146 5452 9696 5023/24
SEEN-BY: 5030/1081 5051/44 5075/21 35 6035/3
@PATH: 5075/35 280/5555 5020/715



   GoldED+ VK   │                                                 │   09:55:30    
                                                                                
В этой области больше нет сообщений.

Остаться здесь
Перейти к списку сообщений
Перейти к списку эх