Researchers introduce byteification for language models to work with characters — Nature
Researchers have described the byteification approach, which adapts token-based large language models to process text at the byte level, Nature reports. This enables such models to work with individual characters in words.
Why models make mistakes
Most large language models encode words as tokens — sequences of letters or parts of words. This approach makes it possible to achieve strong results in many tasks, but it does not provide direct access to individual characters encoded in binary sequences — bytes.
More current news is available on the UA.News Telegram channel Telegram.
As a result, models can make mistakes when counting letters. For example, the letter r occurs three times in the word strawberry, although many large language models answer that it occurs twice.
Working at the byte level
As Nature notes, Minixhofer and co-authors presented byteification as a way to refine existing tokenized models for work at the byte level. The authors showed that models after such refinement can demonstrate competitive results while retaining the ability to read individual characters.