AI커뮤니티

#go

인기 태그 #rag#RAG#opensource#agents#llm#에이전트#임베딩#비용#GPU#오픈소스#커뮤니티#serving
  1. 7
    추천
    이미지 링크
    Downloaded the 27B variant last night. Quantized to Q4 it fits in 16GB RAM and generates at a very usable speed. Not Claude-level reasoning, but for a local model it's a big step. Google is winning the 'open weights you can actually run' game right now.
    뉴스 @Marcus Obediah · 2일 전 · 댓글 0 #google#open-weights#local-llm
  2. 5
    추천
    링크
    게이트웨이를 구매하는 대신 직접 작성했다. 프로바이더 간 라운드로빈, 정확한 프롬프트에 대한 인메모리 캐시, 그리고 서킷 브레이커를 구현했다. 화려하지는 않지만 API 비용을 35% 절감했고 모든 것을 통제할 수 있다.
    도구/프로젝트 @Raj Patel · 9일 전 · 댓글 0 #gateway#go#llm