CodeBERTScore: an automatic metric for code generation, based on BERTScore
-
Updated
Mar 1, 2024 - Jupyter Notebook
CodeBERTScore: an automatic metric for code generation, based on BERTScore
EMNLP'2022: BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation
LLM Evaluation and Observability System for Football Content
MAchine Translation Evaluation Online (MATEO)
Domain-specialized Gemma 2 27B + Gemma 4 31B for SEC filings — fine-tuned on TPU v6e-8 with PyTorch/XLA FSDPv2, plus a Vertex AI Vector Search RAG demo (69 tickers × 381 filings). Same LoRA recipe, +3.5% / +5.8% BERTScore F1.
Fine-tuning GPT-3.5 and Llama3 LLMs for enhanced persona consistency in chatbots using Google's Synthetic Persona Chat dataset
ViAG: A Novel Framework for Fine-tuning Answer Generation models ultilizing Encoder-Decoder and Decoder-only Transformers's architecture
Agent Orchestration - LLM for Legal Metadata Extraction: A Comparative Analysis of Efficiency and Precision (paper 161 PROPOR)
Medical Question Answering System using on PubMed dataset.
The work presented was developed during the internship, as researchers in the field of Natural Language Generation, at the Insid&s Lab laboratory in Milan-Bicocca. The work carried out deals with the creation of a framework for the correct assessment of the impact of the quality of the input datasets on the quality of the text generated by the N…
Implementation of a task-specific QLoRA supervised fine-tuning pipeline for LLaMA-2-7B-Chat, developed for an independent study on structured cover letter generation.
An LLM-powered application that summarizes scientific research papers, extracts tables, and provides real-time evaluation using BERTScore (F1) and ROUGE. Built using Meta’s LLaMA 3–8B via Groq, with table extraction powered by pdfplumber and pandas.
AI-powered research paper summarizer built with Streamlit and Google Gemini. Supports customizable summaries, paper Q&A, ROUGE/BERTScore evaluation, and latency benchmarking.
Role-aware crisis tweet summarization using transformer baselines, automated reward scoring, and preference optimization.
[radscore] Evaluation metrics for radiology report generation — BLEU, ROUGE, BERTScore, F1-RadGraph, CheXbert F1 & GREEN in one `pip install`.
Model-agnostic toolkit to evaluate text-summarization models — ROUGE/BLEU/BERTScore + perplexity + LLM-as-judge, MLflow tracking, CLI + library with unit-tested metrics.
Detecting and Mitigating Self Contradictory Hallucinations in LLMs using a Multi-Agent System and Stepback Prompting
Evaluation framework for measuring LLM reasoning and translation performance across low-resource languages using glossary-guided prompting and semantic similarity metrics.
Add a description, image, and links to the bertscore topic page so that developers can more easily learn about it.
To associate your repository with the bertscore topic, visit your repo's landing page and select "manage topics."