SShortSingh.
Back to feed

Falsify — Find the Smallest Input That Breaks Your C++ Solution (Local AI, No API Keys)

0
·7 views

What I Built Every competitive programmer knows the pain of passing all sample cases but getting a "Wrong Answer" on submission. The classic fix is stress testing, which means writing three separate programs every time you're stuck. I built this for a friend (and myself) who spends more time setting up stress tests than actually solving problems. Demo bash ollama pull gemma4:latest git clone https://github.com/Sohith2007/falsify.git http://localhost:8000 to use the tool. Here is a look at the Web UI finding a deliberate integer overflow bug in real-time: [UPLOAD IMAGE 1 HERE: ui_full_page_1791

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Sidekick part 4: safety as honestly labeled ceilings

Sidekick Part 4: safety as a stack of ceilings, each honestly labeled Most agent safety I've seen is a paragraph in the system prompt: "be careful with destructive commands." Small models ignore system prompts — we established that in Part 2. So Sidekick's safety model assumes the model will disobey and enforces the boundaries in code instead. This is Part 4: approvals, hard refusals, egress control, and the audit ledger — plus the ceilings, stated rather than hidden. Reads auto-run. Everything else sits in APPROVAL_TOOLS — writes, deletes, general shell — gated behind inline [y/N] prompts, wi

0
ProgrammingDEV Community ·

I built my friend a Spanish tutor that runs on her laptop and never phones home

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend I built a Spanish practice partner for a friend who is learning from scratch and freezes when a real person talks to her. Ordering food is the specific thing: she can read a menu, but when the waiter asks "¿algo más?" she has to stop and translate in her head, and the moment passes. So I made LingoBuddy — a terminal program she talks to instead. It holds a conversation with her in Spanish at her level, translates every line into English so she can check her own meaning, and corrects her when she drops an accent. I

0
ProgrammingDEV Community ·

Tuning a Local Qwen Coding Agent: What Changed, What Still Fails

On eight development coding cases repeated three times, a local Qwen3.8-27B agent went from 5/24 to 23/24 functional and delivered successes across sequential output-budget and reasoning rounds. That is promising development evidence. It does not establish generalization or show that xhigh alone caused the difference. The next experiment was less encouraging: lowering temperature to 0.8 on two selected difficult cases preserved 5/6 functional successes but reduced delivered successes to 4/6. The reused temperature-1.0 controls scored 5/6 on both metrics.