“Cline-style agents re-send large conversation histories every step, so most of your tokens are input, not output. Shrink what gets resent (sub-agent scoping, compact modes, tight repo context) — that's where most "1B tokens for $20" bills actually come from.”
“At this rate, my main tasks might take over billion tokens and that is roughly $20 or more. I know the cost calculation is not exactly as simple just like that but is there any other harness or agent I can try that might provide more efficiency or less expense?”
“I subscribed to Cline Pass, and it has been working great for my vibe-coding work. However, I’ve been trying to use a model from Cline Pass with my Hermes agent. It works fine for normal chats, but it doesn’t seem to work with CRON jobs.”
“Any thoughts on why Cline is insisting on authing even though I'm running a local OpenAPI compatible endpoint? Is Cline no longer using the IntelliJ settings for the various endpoints, models, etc.?”