mflux-manual-testing
Provides a change-driven checklist to verify CLI functionality and image outputs for mflux models.
Install
mkdir -p .claude/skills/mflux-manual-testing && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2347" && unzip -o skill.zip -d .claude/skills/mflux-manual-testing && rm skill.zipInstalls to .claude/skills/mflux-manual-testing
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Manually validate mflux CLIs by exercising the changed paths and reviewing output images/artifacts.Key capabilities
- →Validate CLI entrypoints and callback behavior
- →Verify model saving and weight loading paths
- →Test stepwise output generation
- →Perform low-RAM generation checks
- →Compare mflux outputs against diffusers reference
How it works
It provides a change-driven checklist to manually exercise specific code paths and visually inspect generated images and metadata for regressions.
Inputs & outputs
When to use mflux-manual-testing
- →Verifying CLI generation changes
- →Testing model loading paths
- →Confirming stepwise output quality
About this skill
mflux manual testing
Some regressions (especially in CLIs and image IO) are easiest to catch by running the commands and visually inspecting outputs. This skill provides a lightweight, change-driven manual test checklist.
When to Use
- You changed any CLI entrypoint(s) under
src/mflux/models/**/cli/. - You touched callbacks (e.g. stepwise output, memory saver) or metadata/image saving.
- Tests are green but you want confidence in real command usage.
Strategy (change-driven)
- Identify what changed on your branch (new flags, default behavior changes, new callbacks, new models).
- Only run manual checks for the touched areas; don’t try to exercise every CLI.
- Prefer 1–2 seeds and a small step count (e.g. 4) for fast iteration, unless the change affects convergence/quality.
- Before manual CLI testing, reinstall the local tool executables so you’re testing the latest code:
uv tool install --force --editable --reinstall .
Core CLI checks (pick what’s relevant)
- Basic generation: run the CLI once with a representative prompt and confirm the output is not “all noise”.
- Model saving (if relevant): if you touched weight loading/saving or model definitions, run
mflux-savefor the affected model(s) and verify:- the output directory is created
- the command completes without missing-file errors
- Run from disk (if relevant): if you touched save/load paths or model resolution, generate from a locally saved model directory by passing
--model /full/path/to/saved-modeland confirm it runs and produces a sane image. - Stepwise outputs (if relevant): run with
--stepwise-image-output-dirand confirm:- step images are written for each step
- the final step image matches the final output image qualitatively
- the composite image is created
- Low-RAM path (if relevant): run with
--low-ramand confirm:- generation completes
- output quality is sane (no unexpected all-noise output)
- Metadata (if relevant): run with
--metadataand confirm the.metadata.jsonsidecar is emitted and looks consistent.
Output review (human-in-the-loop)
- Always point the human reviewer at:
- the final output image path
- any stepwise directory / composites
- any metadata JSON files
- Ask the human to visually confirm “looks correct” rather than attempting pixel-perfect parity manually.
diffusers reference comparison (new model ports)
mflux does not install diffusers; use a sibling clone (commonly ../diffusers on Desktop).
When: validating a new MLX port before merge, or when golden tests / visuals look wrong.
Match settings: same prompt, width, height, seed, steps, guidance. For fair speed comparisons, match precision (mflux bf16 vs diffusers bf16, not -q 8 vs bf16).
diffusers setup tips (read the reference pipeline first):
- Compare
model_index.json/from_pretrainedkwargs with mfluxget_download_patterns()— disable or passNonefor components mflux does not load local_files_only=True/HF_HUB_OFFLINE=1to use cache and surface missing files early- Distilled vs base variants are often the same pipeline class with different checkpoint + steps/guidance
Run both:
# mflux
uv run mflux-generate-<model> --prompt "..." --width 640 --height 368 --seed 7 --steps 8 --guidance 1.0 --output /tmp/mflux.png
# diffusers (from diffusers repo)
cd ../diffusers && HF_HUB_OFFLINE=1 uv run python -c "..." # inline script; see mflux-debugging
If visuals differ: do not assume the port is broken. Run the latent injection workflow in mflux-debugging to separate (a) transformer/VAE quality from (b) RNG/scheduler differences.
Save comparison PNGs to explicit paths; report /usr/bin/time -p totals. Do not commit comparison artifacts.
Notes
- If the installed
uv toolexecutable behaves differently fromuv run python -m ..., prefer the local module run to isolate environment/tooling issues. - If you need to reinstall the local tool executables, see the repo rules for the current recommended command.
When not to use it
- →Automated unit testing
- →Performance benchmarking without visual validation
Prerequisites
Limitations
- →Manual process requires human review
- →Dependent on local environment setup
How it compares
It prioritizes human-in-the-loop visual verification over pixel-perfect automated testing for CLI and IO changes.
Compared to similar skills
mflux-manual-testing side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| mflux-manual-testing (this skill) | 2 | 2mo | Review | Beginner |
| python-testing-patterns | 77 | 2mo | Review | Intermediate |
| python-playground | 2 | 4mo | Review | Beginner |
| examples-auto-run | 2 | 3mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by filipstrand
View all by filipstrand →You might also like
python-testing-patterns
wshobson
Implement comprehensive testing strategies with pytest, fixtures, mocking, and test-driven development. Use when writing Python tests, setting up test suites, or implementing testing best practices.
python-playground
pydantic
Run and test Python code in a dedicated playground directory. Use when you need to execute Python scripts, test code snippets, investigate CPython behavior, or experiment with Python without affecting the main codebase.
examples-auto-run
openai
Run python examples in auto mode with logging, rerun helpers, and background control.
extract-fuzzer-repro
noir-lang
Extract a Noir reproduction project from fuzzer failure logs in GitHub Actions. Use when a CI fuzzer test fails and you need to create a local reproduction.
klingai-known-pitfalls
jeremylongshore
Manage avoid common mistakes when using Kling AI. Use when troubleshooting issues or learning best practices to prevent problems. Trigger with phrases like 'klingai pitfalls', 'kling ai mistakes', 'klingai gotchas', 'klingai best practices'.
netalertx-testing-workflow
netalertx
Run and debug tests in the NetAlertX devcontainer. Use this when asked to run tests, check test failures, debug failing tests, or execute pytest.