this post was submitted on 14 Mar 2024
68 points (91.5% liked)

Technology

58315 readers
4916 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 1 year ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] kevincox@lemmy.ml 6 points 6 months ago

Absolutely. They are sort of a compression scheme so the tokens contain different numbers of characters based on how frequent that string is. So common words like "the" will typically be one token, or maybe even common phrases like "I am". On the other hand rare punctuation such as "~" may be its own token. There will also be tokens for many common prefixes and suffixes such as "non" and "n't". The tokens of each model are different but they definitely vary in length.