Skip to content

Latest commit

 

History

History
191 lines (143 loc) · 13.4 KB

File metadata and controls

191 lines (143 loc) · 13.4 KB

Prompt Optimization

Overview

Prompt optimization in LiSSA-RATLR enables the automatic systematic refinement of prompts used for traceability link recovery. By leveraging various optimization strategies and evaluation metrics, the effectiveness of prompts may be increased, leading to improved classification accuracy and overall performance. This also enables us to quantify the importance of well designed prompts in the context of traceability link recovery.

Core Components

Overview of Prompt Optimization Subcomponents

The table below provides a brief overview of the subcomponents used in the prompt optimization module.

Component SampleStrategy Selector Metric
Location promptoptimizer.samplestrategy promptoptimizer.promptselector promptoptimizer.promptmetric
Purpose Select items from a collection Orchestrate prompt evaluation with budget Calculate performance scores
Answers "Which generic items to use?" "Which prompts to test when?" "How good is this prompt?"
Method sample(items, sampleSize) selectAndEvaluate(prompts, examples, metric) getMetric(prompts, examples)
Algorithm Examples First/Ordered/Shuffled Simple/UCB Bandit Pointwise/FBeta

Sample Strategies (samplestrategy package)

A SampleStrategy determines how to select a subset of items from a collection. These strategies are used throughout the optimization process to sample items when the full set would be too large or expensive to process. The key method sample(items, sampleSize) returns a list of selected items based on the strategy's selection logic. In practice items may be classification examples, candidate prompts, or simple identifiers depending on the context in which the sampler is used.

Custom sample strategies can be added by implementing the SampleStrategy interface and integrating them via the static factory method SampleStrategy.createSampler(...) defined there.

Available Sample Strategies

  • First Sampler (first): Selects the first n items from the collection without any modification. Maintains the original order of items.
  • Ordered First Sampler (ordered): Sorts items before selecting the first n items. Ensures deterministic sampling based on the natural ordering of items.
  • Shuffled First Sampler (shuffled): Randomly shuffles items before selecting the first n items. Provides random sampling with reproducibility through seeded random number generation.

Prompt Metrics (promptmetric package)

A Metric is a numeric measure used to evaluate the quality of prompts during the optimization process. They are used to guide the optimization by providing feedback on how well a prompt performs in generating accurate traceability links. Currently, they are divided into two types of metrics. Global metrics evaluate the prompt's performance across the entire test dataset. Pointwise metrics scores the performance of prompts on individual data points and reduces the results into a single numeric performance value. If a pointwise metric is used, different scoring and reduction strategies can be configured and combined as desired.

Custom metrics can be added either through implementation of the Global Metrics abstract class or through implementing new scoring and reduction strategies for pointwise metrics.

Available Metrics

  • Global Metrics:
    • F_Beta-Score (fBeta or f1)
  • Pointwise Metrics (pointwise):
    • Scoring Strategies:
      • Binary Scorer (Correct Classification / Incorrect Classification)
    • Reduction Strategies:
      • Mean
  • Mock Metric (mock): Returns dummy values for testing purposes

Selectors (promptselector package)

A Selector orchestrates the evaluation of multiple prompts within a given evaluation budget. They determine which prompts to test and when, managing the trade-off between exploration (testing new prompts) and exploitation (focusing on promising prompts). Selectors use the selectAndEvaluate method to coordinate prompt evaluation, calling the metric to score prompts against classification examples while respecting budget constraints.

The exact evaluation budget parameters are selector-specific, controlling how many total evaluations can be performed. This budget management is crucial for expensive LLM-based evaluations.

Custom selectors can be added by implementing the Selector interface.

Available Selectors

  • Simple Selector (simple or bruteforce): Evaluates all provided candidate prompts against a subset of examples. The sample size is determined by dividing the evaluation budget by the number of prompts. Examples are shuffled randomly before selection to ensure diverse evaluation.

  • Upper Confidence Bound Bandit Selector (ucb): Implements a multi-armed bandit approach using the UCB (Upper Confidence Bound) algorithm. Balances exploration and exploitation by selecting prompts based on both their current performance and uncertainty. More efficient than simple selection when evaluating many prompts, as it focuses on promising candidates.

Optimizers (promptoptimizer package)

The Optimizer module handles prompt optimization requests. Different optimization strategies are implemented to improve prompts using various means. Optimization approaches will usually utilize an iterative process. Prompts are refined over multiple iterations based on the feedback provided through the selected prompt metric. They are highly configurable with the optimization configuration file.

Prompt optimizers reuse the front half of the evaluation pipeline — artifact loading, preprocessing, embedding creation and element stores — plus the configured classifier, aggregator and postprocessor as the metric's scoring machinery. They do not run the evaluation's own classification, aggregation and statistics stages, and they write results-prompt-optimization-*.md rather than results-*.md and traceLinks-*.csv. Both halves share LiSSA's caching mechanism, so repeated optimizer runs are consistent and reproducible.

Custom optimizers can be added by implementing the Prompt Optimizer interface.

Available Optimizers

  • Naive Iterative Optimizer (iterative or simple): The most basic optimizer that makes changes to the prompt in each iteration. It simply queries the large language model to improve the current prompt using an optimization prompt. The new prompt is naively carried over to the next iteration without any further checks.

    • simple: Defaults to one (1) iteration
    • iterative: Defaults to five (5) iterations
  • Feedback-Based Optimizer (feedback): The iterative feedback optimizer improves prompts by leveraging feedback from the large language model. In each iteration, it queries the model with an additional feedback text on the current prompt. The optimizer carries the optimized prompt to the next iteration naively. Trace links that were incorrectly classified in previous iterations are highlighted in the feedback text to guide the model towards better performance.

  • ProTeGi Optimizer (protegi or gradient): An advanced optimizer based on textual gradient descent for large language models, following the approach by Pryzant et al. (2023). Uses textual gradients derived from error analysis to systematically refine prompts. In each iteration:

    1. Candidate Expansion: Generates multiple candidate prompt variations
      • Analyzes why the current prompt misclassifies examples (textual gradients)
      • Creates transformations based on these error patterns
      • Generates synonym variations to explore the prompt space
    2. Candidate Evaluation: Uses the configured selector and metric to evaluate all candidate prompts
      • Selector decides which candidate prompts to test and on how many examples (budget-aware)
      • Metric scores each candidate prompt's performance
    3. Best Selection: Selects the top-performing candidate prompts (beam size) for the next iteration

    Example flow: Current prompt gets accuracy 70% → generates 20 candidates → evaluates them with limited budget → selects top 4 for next iteration

  • Mock Optimizer (mock): Returns dummy optimized prompts for testing purposes

Configuration

Optimization Configuration Structure

Modules of the evaluation configuration file will also need to be configured in the optimization configuration file. This excerpt shows the additional configuration options specific to prompt optimization.

{
  "metric": {
    "name": "mock",
    "args": {}
  },
  "selector": {
    "name": "ucb",
    "args": {
      "samples_per_eval": 16
    }
  },
  "prompt_optimizer": {
    "name": "simple_openai",
    "args": {
      "prompt": "Question: Here are two parts of software development artifacts.\n\n            {source_type}: '''{source_content}'''\n\n            {target_type}: '''{target_content}'''\n            Are they related?\n\n            Answer with 'yes' or 'no'.",
      "model": "gpt-4o-mini-2024-07-18"
    }
  }
}

These three keys are added to the modules of a regular evaluation configuration; the snippet above shows only the additions.

selector is optional in general but required for the protegi / gradient optimizer — omitting it aborts the run with Selector must not be null for ProTeGi optimizers.

The ProTeGi optimizer accepts a large set of arguments; the ones most often set are maximum_iterations, minibatch_size (default 64), beam_size (default 4), max_expansion_factor (default 8) and number_of_gradients (default 4). See example-configs/gradient-optimizer-config.json for a working example.

To see detailed configurable fields for any of the modules refer to a prompt optimization result file. After executing a minimal configuration the resulting file will contain the full configuration with all default values filled in.

Note

An optimization configuration is a different schema from an evaluation configuration: it is accepted by lissa optimize and rejected by lissa eval. See the Configuration Guide.

Usage

Refer to the CLI Documentation for instructions on how to run prompt optimization using the command line interface.

Optimization Process

The optimization process generally follows these steps:

  1. Baseline Evaluation (Optional): If evaluation configurations are provided, the baseline performance of the original prompt is measured.
  2. Prompt Optimization: The prompt optimizer is executed using the specified optimization configuration. The prompt is refined iteratively based on the selected metric.
  3. Post-Optimization Evaluation (Optional): If evaluation configurations are provided, the optimized prompt is evaluated to measure differences over the baseline.

Output and Results

Result Files

The prompt optimization results will be stored as results-prompt-optimization-<config-file-name>_<uuid>.md in the current working directory, just as regular evaluation results. The <uuid> is a hash over the fully-resolved configuration, so re-running the same configuration overwrites the previous file. They include the full configuration used for optimization as well as the optimized prompt.