GLOSSIKON
โ† Back to blog

Glossikon's Vector Database

So Glossikon leans quite heavily into LLMs. It's for makeup languages, LLM literally has the word language in the name. Infact the original idea for Glossikon was, now that we have LLMs, let's make something todo with languages. Dare I? Could I? Maybe, make my own language? No-one on earth would be able to understand, except me and my bot. Actually not even me! Bots are going to argue with each other in their own language! I have heard about bots developing their own shorthand with each other bit this is taking it to another level. I love it.

Just one problem though- when a language gets too big, like 1000+ words I ran out of context of my local LLM running on my 5070ti (BTF edition mind you ๐Ÿ˜‰) at 10k words it overflows even the biggest models in the world!

Of course my first thought was I probly needed a vector database right? That's the ol' go to for reducing context size in my mind. What I was seeing was only the last two phases where too big, the lexicon and the semantic meaning. The syntax and the symbols used were never that big.

Brain wave! That sounds like a dictionary to me! A lexicon with semantic meaning ๐Ÿ˜Œ so I separated out a dictionary and plopped it into a vector database. When translating a sentence, Glossikon looks up words in the dictionary by semantic meaning so synonyms are also returned. And finally Glossikon MCP tries it all together and allows the local LLM to use the dictionary during translations. Phew ๐Ÿ˜