Prompt optimization in LiSSA-RATLR enables the automatic systematic refinement of prompts used for traceability link recovery. By leveraging various optimization strategies and evaluation metrics, the effectiveness of prompts may be increased, leading to improved classification accuracy and overall performance. This also enables us to quantify the importance of well designed prompts in the context of traceability link recovery.
The table below provides a brief overview of the subcomponents used in the prompt optimization module.
| Component | SampleStrategy | Selector | Metric |
|---|---|---|---|
| Location | promptoptimizer.samplestrategy |
promptoptimizer.promptselector |
promptoptimizer.promptmetric |
| Purpose | Select items from a collection | Orchestrate prompt evaluation with budget | Calculate performance scores |
| Answers | "Which generic items to use?" | "Which prompts to test when?" | "How good is this prompt?" |
| Method | sample(items, sampleSize) |
selectAndEvaluate(prompts, examples, metric) |
getMetric(prompts, examples) |
| Algorithm Examples | First/Ordered/Shuffled | Simple/UCB Bandit | Pointwise/FBeta |
A SampleStrategy determines how to select a subset of items from a collection.
These strategies are used throughout the optimization process to sample items when the full set would be too large or expensive to process.
The key method sample(items, sampleSize) returns a list of selected items based on the strategy's selection logic.
In practice items may be classification examples, candidate prompts, or simple identifiers depending on the context in which the sampler is used.
Custom sample strategies can be added by implementing the SampleStrategy interface and integrating them via the static factory method SampleStrategy.createSampler(...) defined there.
First Sampler(first): Selects the first n items from the collection without any modification. Maintains the original order of items.Ordered First Sampler(ordered): Sorts items before selecting the first n items. Ensures deterministic sampling based on the natural ordering of items.Shuffled First Sampler(shuffled): Randomly shuffles items before selecting the first n items. Provides random sampling with reproducibility through seeded random number generation.
A Metric is a numeric measure used to evaluate the quality of prompts during the optimization process.
They are used to guide the optimization by providing feedback on how well a prompt performs in generating accurate traceability links.
Currently, they are divided into two types of metrics.
Global metrics evaluate the prompt's performance across the entire test dataset.
Pointwise metrics scores the performance of prompts on individual data points and reduces the results into a single numeric performance value.
If a pointwise metric is used, different scoring and reduction strategies can be configured and combined as desired.
Custom metrics can be added either through implementation of the Global Metrics abstract class or through implementing new scoring and reduction strategies for pointwise metrics.
Global Metrics:- F_Beta-Score (
fBetaorf1)
- F_Beta-Score (
Pointwise Metrics(pointwise):- Scoring Strategies:
- Binary Scorer (Correct Classification / Incorrect Classification)
- Reduction Strategies:
- Mean
- Scoring Strategies:
Mock Metric(mock): Returns dummy values for testing purposes
A Selector orchestrates the evaluation of multiple prompts within a given evaluation budget.
They determine which prompts to test and when, managing the trade-off between exploration (testing new prompts) and exploitation (focusing on promising prompts).
Selectors use the selectAndEvaluate method to coordinate prompt evaluation, calling the metric to score prompts against classification examples while respecting budget constraints.
The exact evaluation budget parameters are selector-specific, controlling how many total evaluations can be performed. This budget management is crucial for expensive LLM-based evaluations.
Custom selectors can be added by implementing the Selector interface.
-
Simple Selector(simpleorbruteforce): Evaluates all provided candidate prompts against a subset of examples. The sample size is determined by dividing the evaluation budget by the number of prompts. Examples are shuffled randomly before selection to ensure diverse evaluation. -
Upper Confidence Bound Bandit Selector(ucb): Implements a multi-armed bandit approach using the UCB (Upper Confidence Bound) algorithm. Balances exploration and exploitation by selecting prompts based on both their current performance and uncertainty. More efficient than simple selection when evaluating many prompts, as it focuses on promising candidates.
The Optimizer module handles prompt optimization requests.
Different optimization strategies are implemented to improve prompts using various means.
Optimization approaches will usually utilize an iterative process.
Prompts are refined over multiple iterations based on the feedback provided through the selected prompt metric.
They are highly configurable with the optimization configuration file.
Prompt optimizers reuse the front half of the evaluation pipeline — artifact loading, preprocessing, embedding creation and element stores — plus the configured classifier, aggregator and postprocessor as the metric's scoring machinery.
They do not run the evaluation's own classification, aggregation and statistics stages, and they write results-prompt-optimization-*.md rather than results-*.md and traceLinks-*.csv.
Both halves share LiSSA's caching mechanism, so repeated optimizer runs are consistent and reproducible.
Custom optimizers can be added by implementing the Prompt Optimizer interface.
-
Naive Iterative Optimizer(iterativeorsimple): The most basic optimizer that makes changes to the prompt in each iteration. It simply queries the large language model to improve the current prompt using an optimization prompt. The new prompt is naively carried over to the next iteration without any further checks.simple: Defaults to one (1) iterationiterative: Defaults to five (5) iterations
-
Feedback-Based Optimizer(feedback): The iterative feedback optimizer improves prompts by leveraging feedback from the large language model. In each iteration, it queries the model with an additional feedback text on the current prompt. The optimizer carries the optimized prompt to the next iteration naively. Trace links that were incorrectly classified in previous iterations are highlighted in the feedback text to guide the model towards better performance. -
ProTeGi Optimizer(protegiorgradient): An advanced optimizer based on textual gradient descent for large language models, following the approach by Pryzant et al. (2023). Uses textual gradients derived from error analysis to systematically refine prompts. In each iteration:- Candidate Expansion: Generates multiple candidate prompt variations
- Analyzes why the current prompt misclassifies examples (textual gradients)
- Creates transformations based on these error patterns
- Generates synonym variations to explore the prompt space
- Candidate Evaluation: Uses the configured selector and metric to evaluate all candidate prompts
- Selector decides which candidate prompts to test and on how many examples (budget-aware)
- Metric scores each candidate prompt's performance
- Best Selection: Selects the top-performing candidate prompts (beam size) for the next iteration
Example flow: Current prompt gets accuracy 70% → generates 20 candidates → evaluates them with limited budget → selects top 4 for next iteration
- Candidate Expansion: Generates multiple candidate prompt variations
-
Mock Optimizer(mock): Returns dummy optimized prompts for testing purposes
Modules of the evaluation configuration file will also need to be configured in the optimization configuration file. This excerpt shows the additional configuration options specific to prompt optimization.
{
"metric": {
"name": "mock",
"args": {}
},
"selector": {
"name": "ucb",
"args": {
"samples_per_eval": 16
}
},
"prompt_optimizer": {
"name": "simple_openai",
"args": {
"prompt": "Question: Here are two parts of software development artifacts.\n\n {source_type}: '''{source_content}'''\n\n {target_type}: '''{target_content}'''\n Are they related?\n\n Answer with 'yes' or 'no'.",
"model": "gpt-4o-mini-2024-07-18"
}
}
}These three keys are added to the modules of a regular evaluation configuration; the snippet above shows only the additions.
selector is optional in general but required for the protegi / gradient optimizer — omitting it aborts the run with Selector must not be null for ProTeGi optimizers.
The ProTeGi optimizer accepts a large set of arguments; the ones most often set are maximum_iterations, minibatch_size (default 64), beam_size (default 4), max_expansion_factor (default 8) and number_of_gradients (default 4). See example-configs/gradient-optimizer-config.json for a working example.
To see detailed configurable fields for any of the modules refer to a prompt optimization result file. After executing a minimal configuration the resulting file will contain the full configuration with all default values filled in.
Note
An optimization configuration is a different schema from an evaluation configuration: it is accepted by lissa optimize and rejected by lissa eval. See the Configuration Guide.
Refer to the CLI Documentation for instructions on how to run prompt optimization using the command line interface.
The optimization process generally follows these steps:
- Baseline Evaluation (Optional): If evaluation configurations are provided, the baseline performance of the original prompt is measured.
- Prompt Optimization: The prompt optimizer is executed using the specified optimization configuration. The prompt is refined iteratively based on the selected metric.
- Post-Optimization Evaluation (Optional): If evaluation configurations are provided, the optimized prompt is evaluated to measure differences over the baseline.
The prompt optimization results will be stored as results-prompt-optimization-<config-file-name>_<uuid>.md in the current working directory, just as regular evaluation results. The <uuid> is a hash over the fully-resolved configuration, so re-running the same configuration overwrites the previous file.
They include the full configuration used for optimization as well as the optimized prompt.