Skip to main content

Indie game storeFree gamesFun gamesHorror games
Game developmentAssetsComics
SalesBundles
Jobs
TagsGame Engines

Hi foxehtroppy! This is possible, but is a lower priority item for me. I spent a lot of time on koboldcpp with local models and don't think any would be smart/fast enough to handle the game's needs. Out of curiosity, what model would you be using with koboldcpp? And would you be willing to test the feature for me if I email you a build with it?

(3 edits) (+1)

Well the model I use is Cydonia-24B-v4.3-absolute-heresy.i1-Q4_K_M, is more in the NSFW side but so far it have been handling well enough and so far the bot doesn't go thirsty if I don't input any lewd stuff.

Speed wise: is more or less slow(1-3 words per second), sometimes go fast(maybe 4-6words per second) for some reason. but I dont mind the slowness.

Smartness wise it appears to handle decently when there are two or three characters (when using a CYOA char bot). I still need to check when using two bots in a group chat. Couple of times it got names the own model generated with the CYOA char bot.

And sure, though I suck in setting the min P and those type of parametes (I exported .json with those parameters already tweaked for Silly Tavern)


*But if it is a low priority it can wait :D

Cydonia-24B has been pretty good in my experience, though I did run into the same issue of it being a little slow, like you said. I will experiment with koboldcpp and add it in a future update if the quality is okay. Thanks!

hey if you need another tester for local Ilm, I got a 5060 to 16gb and so far I run 24 well in that area with fast speeds and even use an agent when I use marinara engine. I think the option of local would be awesome!

Hey ElPapaDon! I added local LLM support in v0.1.2. If you try it out, please let me know how it goes and what models work well

gotcha! ill give some feedback as i have today off so ill be touching upon this game alot xD

(1 edit)

What GPU are you using? I'm using the exact same model and quant, It's taking forever for me to load it up via llama.cpp on an RX 7900 XTX and I keep getting an error that it lost connection to the endpoint was lost mid-reply, is there something specific you did, what context size are you using?