Im trying to convert image based tutorials into text based markdown. The images themselves should be much less than 1000 tokens. I already cant copy/paste with Leo back and forth between the terminal. Been documented on the github for months. I try to convert image based tutorials for a big development into a text with Sonnet. Each image is probably less than 500 tokens. I process 37 images across 2 conversations. 3rd conversation Im greeted by forced Haiku downgrading. Its conversion and formatting is vastly inferior to Sonnets. It didnt even understand the primer to start the image to text extraction process.
My service has been degraded for that past 2-3 hours.
Im assuming the copy/paste response is also a response to my other passive/aggressive posts or comments that are not explicitly inflammatory but critical of the platform that keep coming back as “post hidden by community flags”. Suppressing deserved criticism, even if it is packaged in a way that is not preferred, is a bad look. The resolution is solving the problem, not suppressing my criticism.
The issue with the image processing in particular is the context consumption had to have been less than 100k tokens of conversation across the entire account for the entire morning before even starting the image processing. Processing just those images destroyed my premium usage for hours. My question is less about being notified about my quota being hit. Its about why images who text conversion would be anywhere between 1000-5000 tokens max, would have eaten nearly 12 hours worth of quota. Its screen snips of 70% of a PDF page of documentation im converting to text only. I need to convert 112 pages. This is less about “When will I hit my quota” and more about “can this tool still fill me daily work flow with its cost/benefits incentive”. I maybe wrong and image processing is a huge resource consumer. Resent tests show image processing has lower token consumption than text.
Weird. Ive been interacting with leo on my phone a lot this morning. Think phone keyboard = throttled input = low token input. There would be no leo memories to invisibly eat up context either. I’ve never hit a full context window on mobile. That would be pretty nuts. Ive had 3 short conversations. Probably 30 turns a piece. I’m degraded to haiku. Thats less than a full sonnet context window. Only time I’ve ever been degraded like this before is when I had Opus doing multiple review passes on like a 120k word context block. Like I was dropping 120k words at a time into a context window with a primer and having opus iterate through it to do context extraction. It was probably like millions of tokens in couple hours. That was a warranted throttle and I would expect Brave to be like “you got your monies worth, use the penny models for a few” and me be like “fair enough”.
Something is wrong or broken. Either quote detection, massive system prompts are bugged and getting sent silently, or something. I’ve been getting regularly throttled to Haiku for what I previously could have done 100 times over and not hit rate limiting.
Oh interesting, we’ve set the bar for a selected model being ‘downgraded’ very high as it’s not something we want to happen often at all. The idea is to deter bad actors but if it significantly affects real users then it’s unintended
One thing to note is that we don’t currently do rate limiting by token count (we should and it’s in our backlog), so - assuming it is rate limiting kicking in and not a bug - the message length & number of images shouldn’t be a factor. I can take a look at the latest logic for downgrading though, we may have missed something
The data for this instance has been polluted, but the next time I run into it I’ll try to pull the text and at least approximate a token count. It would have been at least 18 hours since I had done the image conversions that caused me to get throttled. Like the image throttling is annoying and I don’t like it, but it’s not something that I can really justify being critical for. It is what it is and I have a a local LLM stack on a 32GB GPU I can run that PDF through, I’m just not super familiar with it yet and didn’t have the epgu setup at the time. I have an alternative.
The throttling from more free form prose interactions across 3 different conversations on mobile in 3 hours was a bit weird though. I was listening to YouTube and planning a couple different projects and brainstorming some stuff. Like conversational notes. Typically shorter exchanges. 3-5 sentance paragraphs. Maybe 10 paragraphs per conversation across all 3 conversations when I realized the last 3 or 4 replies had come from Haiku. I wouldn’t have expected the image conversion dings to have trailed into that quota space over 12 plus hours later.
I have two brave leo subscriptions. One for work and one for personal. The work one has been throttled just that one time where I had Opus running a large primer over a 120k word multi-chat trouble shoot to generate documentation from it. I had to do it multiple times to weight things differently against each other. It was a very expansive canbus reverse engineering project with proprietary PGNs and frame hold and burst logic. That hit the throttle and rightfully so. I must have iterated over that 120k word context like 6 times in 30 minutes. I would expect getting throttled.
The images and chitchat was on my personal. Ive never been throttled on my personal since I’ve had it.
The images was unexpected but I still kinda get it, and I can offload.
Chit-chatting with sonnet over a mobile keyboard was turbo unexpected.
Okay, so I just got throttle on my work account and I have exact stats. Something is extremely broken. Ive been on PTO for 2 days not using my company laptop.
Single conversation
35,168 visible characters
Memories approximately 1.5k tokens
27 total turns
I can’t tell if I can share images or not so I grabbed a proton link. I had haiku do a summary. This brutal lol. It was like 12 large terminal commands and me sharing the outputs, and some prose. This account hasn’t been used in days.
Can I request some sort of partial refund or something? This is brutal. I’ll provide censored screenshot of the volume of interactions I have had only just today. 4 small-medium conversations. Not a single one of them enough to hit a full 200k word context window. Not even enough fo get truncated or compacted or to start choking on context. I am throttled to Haiku. It has been like this for the past 2 weeks across 2 different accounts. The only thing saving the product is it being locally stored and encrypted. Months long copy/paste bug and brutal throttling. Sonnet use in Jetbrains isn’t this extreme and Jetbrains robs you on token use.
Its sonnet. I’m being throttled from just sonnet. If I try to use Sonnet, I get sent to Haiku. I can use Mistral Large and Opus. Ain’t nobody can use Opus. It can’t do anything iterative. It wants to talk about 900 different bugs on the tree when you are trying to find a forest. You fight with the model more than use it. Is this a bug? Thats not a joke.
P. S. Ia per model throttling intended? Like I can switch to Opus, a more expensive model, send large context windows, but can’t select sonnet under the response window to regenerate it with Sonnet, which tells me it’s an intended hard limit per model or a VERY coincidental bug. It would be weird to allow me to send enormous context to Opus(which is a practically useless model as I believe it was patterned after lunatic with neurotic paranoia) but not keep using the cheaper sonnet model.
I need some clarity on this because I pay for multiple annual licenses and tell other people to use the product.