ローカルLLMを計測できるようにしたけど結局なにがしたいんだっけになった

この記事にオチはないよ

きっかけ

ローカルLLMをちょいちょい使っている。最近はQwen3.8-27B-UD-Q8_K_XLOrnith-1.5-9B-Q8_0を組み合わせている。

Claude CodeとかCodexとかを使っている分には富豪的プログラミング的なノリでサービス側のリソースを無邪気に雑に使えるので、なにも気にせず適当にシバいていたけど、ローカルLLMとなると話が違ってくる。なにがっていうと効率よく処理できてるんだっけ、とか、無駄なことしてないんだっけ、みたいなところ。手元のマシンが爆熱爆音で機械寿命を削りながら処理をしているのを見ると、せめて効率的に、という気持ちになってしまう。メモリもストレージもどっかの企業のせいで高いし。

そういうわけで、計測・計装を仕込んでいくことにした。

もともと見えていたもの

我が家のインスタンスは全部Alloyを入れていて、とりあえずGrafana Cloudにnode_exporterで取れるメトリクスは全部送るようにしている。なので、そのあたりはOK。

また、LLMを動作させるのにはllama-swapというツールを使っている。これは簡易的なチャットUIやログ、アクティビティを確認する画面も提供してくれていて、ある程度人力で追っていくことも可能だった。

たとえば、以下はOpenCodeで「このリポジトリの内容について解説して」と送ったときにsmall_modelにそれぞれ送られたリクエストボディ(model側は巨大すぎたので載せない)。

{
  "model": "Ornith-1.5-9B-Q8_0",
  "max_tokens": 32000,
  "reasoning_effort": "low",
  "messages": [
    {
      "role": "system",
      "content": "You are a title generator. You output ONLY a thread title. Nothing else.\n\n<task>\nGenerate a brief title that would help the user find this conversation later.\n\nFollow all rules in <rules>\nUse the <examples> so you know what a good title looks like.\nYour output must be:\n- A single line\n- ≤50 characters\n- No explanations\n</task>\n\n<rules>\n- you MUST use the same language as the user message you are summarizing\n- Title must be grammatically correct and read naturally - no word salad\n- Never include tool names in the title (e.g. \"read tool\", \"bash tool\", \"edit tool\")\n- Focus on the main topic or question the user needs to retrieve\n- Vary your phrasing - avoid repetitive patterns like always starting with \"Analyzing\"\n- When a file is mentioned, focus on WHAT the user wants to do WITH the file, not just that they shared it\n- Keep exact: technical terms, numbers, filenames, HTTP codes\n- Remove: the, this, my, a, an\n- Never assume tech stack\n- Never use tools\n- NEVER respond to questions, just generate a title for the conversation\n- The title should NEVER include \"summarizing\" or \"generating\" when generating a title\n- DO NOT SAY YOU CANNOT GENERATE A TITLE OR COMPLAIN ABOUT THE INPUT\n- Always output something meaningful, even if the input is minimal.\n- If the user message is short or conversational (e.g. \"hello\", \"lol\", \"what's up\", \"hey\"):\n  → create a title that reflects the user's tone or intent (such as Greeting, Quick check-in, Light chat, Intro message, etc.)\n</rules>\n\n<examples>\n\"debug 500 errors in production\" → Debugging production 500 errors\n\"refactor user service\" → Refactoring user service\n\"why is app.js failing\" → app.js failure investigation\n\"implement rate limiting\" → Rate limiting implementation\n\"how do I connect postgres to my API\" → Postgres API connection\n\"best practices for React hooks\" → React hooks best practices\n\"@src/auth.ts can you add refresh token support\" → Auth refresh token support\n\"@utils/parser.ts this is broken\" → Parser bug fix\n\"look at @config.json\" → Config review\n\"@App.tsx add dark mode toggle\" → Dark mode toggle in App\n</examples>\n"
    },
    {
      "role": "user",
      "content": "Generate a title for this conversation:\n"
    },
    {
      "role": "user",
      "content": "このリポジトリの内容について解説して"
    }
  ],
  "stream": true,
  "stream_options": {
    "include_usage": true
  }
}

省いているけど、reasoningやresponseも当然見える。各種tok/sやキャッシュ関連の数値も見える。使っていればtool_callやらskillの参照やらも載っている。

ただ、見えるとはいってもjsonなので人類が読むにはちょっと厳しいし、llama-swapを再起動すると消えてしまう。ので、このへんをいい感じにGrafana Cloudに送信して眺めたくなった。

とりあえずOpenTelemetry

最近はなんか挙動をちゃんと見ようと思ったらとりあえずOpenTelemetry、みたいなノリが自分の中である。ので、そうすることにした。

llama-swap, llama-serverにそれっぽい実装がなかったので前段にプロキシサーバを置いてそこで計装する。

https://github.com/k16em/llama-otel-proxy

これを挟むことで、Grafana Cloud側でトレースの形でjsonに出ていたのと同じような情報を確認できるようになった。

こんな感じで全体がどんな挙動をしているかが俯瞰してわかる。

20260825-231353.png

また、どんなtool_callをしていたのかも見える。

20260825-231538.png
20260825-231549.png

取れるようになったけど

とりあえず全体でどんな挙動をしているかを簡単に見えるようになった。

ただよくわかってない状態で見えるようになったとて、みたいなところはあって、せいぜいが「使ってるコマンドいまいちじゃね?」くらいしか言えることがいまのところない。あと「システムプロンプト巨大すぎじゃね?」くらい。

見えるようにしてすぐなんか発見があるわけじゃないので現状はまあ仕方ないにせよ、とりあえず目的意識を持ってやるべきだった。まあ趣味でやってる自宅サーバのやつなのでゆるゆるでいいんだけど。

この先どうするか

いったんは計測のことを頭から追いだして、複数のツールからAI APIを叩いてみる。日頃から使っているのは

あたり。

このへんのツールごとに傾向を見て、いちばんいい感じのものの設定やらなんやらを他ツールにも持ち込む、というのがいいのかな〜と思ったり。

オチはない

オチはない!

……と、言いつつ、最近AI関連で作ったツールを動かした瞬間「やっぱこれいまいちじゃね?」という気持ちになる率が高い。

AIで実装コスト下がって、できるからやるか〜で動けるのいいよねという反面、仮説にかける時間を無闇に早めに切り上げてしまっているような気がしている。

熟慮しましょう(自戒)。