Un presupuesto de errores contesta la pregunta que ninguna alerta por umbral resuelve: ¿cuántos fallos puedes permitirte antes de que la fiabilidad deje de ser aceptable? En vez de debatir si un 94,2 % frente a un objetivo del 95 % es "grave", fijas un SLO, lo traduces a un número de fallos tolerables y dejas que ese saldo decida cuándo frenar. Lo implementamos aquí en Python y JavaScript.
Qué mide un presupuesto de errores
Cuatro conceptos sostienen el modelo:
- SLO: la tasa de éxito objetivo (por ejemplo, 95 %).
- Presupuesto de errores: la tasa de fallo tolerable; con un SLO del 95 %, puede fallar el 5 %.
- Tasa de consumo (burn rate): con qué rapidez gastas el saldo; 2× lo agota en medio período.
- Ventana: el período de medición, normalmente móvil de 24 horas o 7 días.
Con un SLO del 95 % en 24 horas y 10.000 resoluciones, tu presupuesto es de 500 fallos: mientras te queden puedes desplegar; al llegar a 500, conviene parar los cambios arriesgados. Así "la fiabilidad bajó un poco" pasa a ser una cifra accionable.
Cuándo actuar según la tasa de consumo
El saldo dice cuánto queda; la tasa de consumo, si llegarás al final de la ventana:
| Tasa de consumo | Qué significa | Acción |
|---|---|---|
| < 1,0 | Consumes más lento de lo esperado | No hace falta actuar |
| 1,0 | En ritmo de agotar el saldo al cerrar la ventana | Vigila de cerca |
| 2,0 | Saldo agotado a mitad de la ventana | Investiga y reduce el ritmo |
| 5,0+ | Consumo acelerado del presupuesto | Pausa las resoluciones no críticas |
Rastreador en Python
Guarda cada resultado en una ventana móvil, calcula el saldo y dispara callbacks al cambiar de estado. solve_with_budget lo envuelve y se detiene cuando el saldo se agota:
import time
import threading
from dataclasses import dataclass, field
from collections import deque
from enum import Enum
API_KEY = "YOUR_API_KEY"
class BudgetStatus(Enum):
HEALTHY = "healthy" # Budget > 50% remaining
WARNING = "warning" # Budget 10-50% remaining
CRITICAL = "critical" # Budget < 10% remaining
EXHAUSTED = "exhausted" # Budget depleted
@dataclass
class SLOConfig:
"""Service Level Objective configuration."""
target_success_rate: float = 0.95 # 95%
window_seconds: int = 86400 # 24 hours
warning_threshold: float = 0.50 # Alert at 50% budget
critical_threshold: float = 0.10 # Alert at 10% budget
@dataclass
class ErrorBudgetEvent:
timestamp: float
success: bool
class ErrorBudgetTracker:
"""Tracks error budget consumption for CAPTCHA solving."""
def __init__(self, config: SLOConfig = SLOConfig()):
self.config = config
self._events: deque[ErrorBudgetEvent] = deque()
self._lock = threading.Lock()
self._callbacks: dict[BudgetStatus, list[callable]] = {
status: [] for status in BudgetStatus
}
self._last_status = BudgetStatus.HEALTHY
def on_status_change(self, status: BudgetStatus, callback: callable):
"""Register a callback for status transitions."""
self._callbacks[status].append(callback)
def record(self, success: bool):
"""Record a solve attempt."""
now = time.monotonic()
event = ErrorBudgetEvent(timestamp=now, success=success)
with self._lock:
self._events.append(event)
self._prune(now)
new_status = self._compute_status()
if new_status != self._last_status:
self._last_status = new_status
for cb in self._callbacks.get(new_status, []):
try:
cb(self.get_report())
except Exception as e:
print(f"[BUDGET] Callback error: {e}")
def _prune(self, now: float):
"""Remove events outside the window."""
cutoff = now - self.config.window_seconds
while self._events and self._events[0].timestamp < cutoff:
self._events.popleft()
def _compute_status(self) -> BudgetStatus:
remaining = self.remaining_fraction
if remaining <= 0:
return BudgetStatus.EXHAUSTED
if remaining < self.config.critical_threshold:
return BudgetStatus.CRITICAL
if remaining < self.config.warning_threshold:
return BudgetStatus.WARNING
return BudgetStatus.HEALTHY
@property
def total_events(self) -> int:
with self._lock:
return len(self._events)
@property
def success_count(self) -> int:
with self._lock:
return sum(1 for e in self._events if e.success)
@property
def failure_count(self) -> int:
with self._lock:
return sum(1 for e in self._events if not e.success)
@property
def current_success_rate(self) -> float:
total = self.total_events
return self.success_count / total if total > 0 else 1.0
@property
def error_budget_total(self) -> float:
"""Total allowed failures in the window."""
total = self.total_events
if total == 0:
return 0
return total * (1 - self.config.target_success_rate)
@property
def error_budget_remaining(self) -> float:
"""Remaining failure allowance."""
return max(0, self.error_budget_total - self.failure_count)
@property
def remaining_fraction(self) -> float:
"""Fraction of error budget remaining (0.0 to 1.0)."""
budget = self.error_budget_total
if budget <= 0:
return 1.0 if self.failure_count == 0 else 0.0
return max(0, self.error_budget_remaining / budget)
@property
def burn_rate(self) -> float:
"""How fast the budget is being consumed (1.0 = normal, 2.0 = 2× faster)."""
total = self.total_events
if total == 0:
return 0.0
expected_failures = total * (1 - self.config.target_success_rate)
if expected_failures == 0:
return 0.0
return self.failure_count / expected_failures
def get_report(self) -> dict:
return {
"status": self._last_status.value,
"slo_target": self.config.target_success_rate,
"current_rate": round(self.current_success_rate, 4),
"total_events": self.total_events,
"successes": self.success_count,
"failures": self.failure_count,
"budget_total": round(self.error_budget_total, 1),
"budget_remaining": round(self.error_budget_remaining, 1),
"budget_remaining_pct": round(self.remaining_fraction * 100, 1),
"burn_rate": round(self.burn_rate, 2),
}
# --- Integration with solver ---
budget = ErrorBudgetTracker(SLOConfig(
target_success_rate=0.95,
window_seconds=3600, # 1-hour window for demo
))
# Register alerts
budget.on_status_change(BudgetStatus.WARNING, lambda r:
print(f"[ALERT] Budget warning: {r['budget_remaining_pct']}% remaining"))
budget.on_status_change(BudgetStatus.CRITICAL, lambda r:
print(f"[ALERT] Budget critical: {r['budget_remaining_pct']}% remaining"))
budget.on_status_change(BudgetStatus.EXHAUSTED, lambda r:
print(f"[ALERT] Budget EXHAUSTED — throttle new requests"))
def solve_with_budget(params: dict) -> str:
"""Solve CAPTCHA while tracking error budget."""
import requests
if budget._last_status == BudgetStatus.EXHAUSTED:
raise RuntimeError("Error budget exhausted — solving paused")
try:
submit_params = {**params, "key": API_KEY, "json": 1}
resp = requests.post(
"https://ocr.captchaai.com/in.php", data=submit_params, timeout=30
).json()
if resp.get("status") != 1:
budget.record(False)
raise RuntimeError(f"Submit: {resp.get('request')}")
task_id = resp["request"]
start = time.monotonic()
while time.monotonic() - start < 180:
time.sleep(5)
poll = requests.get("https://ocr.captchaai.com/res.php", params={
"key": API_KEY, "action": "get", "id": task_id, "json": 1,
}, timeout=15).json()
if poll.get("request") == "CAPCHA_NOT_READY":
continue
if poll.get("status") == 1:
budget.record(True)
return poll["request"]
budget.record(False)
raise RuntimeError(f"Solve: {poll.get('request')}")
budget.record(False)
raise RuntimeError("Timeout")
except Exception:
budget.record(False)
raise
# Usage
for i in range(100):
try:
token = solve_with_budget({
"method": "turnstile",
"sitekey": "0x4XXXXXXXXXXXXXXXXX",
"pageurl": "https://example.com",
})
except RuntimeError as e:
if "exhausted" in str(e):
print(f"Stopped at iteration {i}")
break
print(budget.get_report())
Rastreador en JavaScript
La misma lógica en Node.js, para un worker que registre cada resolución:
class ErrorBudgetTracker {
#events = [];
#config;
#callbacks = {};
constructor(config = {}) {
this.#config = {
targetRate: config.targetRate || 0.95,
windowMs: config.windowMs || 3600_000,
warningThreshold: config.warningThreshold || 0.5,
criticalThreshold: config.criticalThreshold || 0.1,
};
this.lastStatus = "healthy";
}
on(status, callback) {
this.#callbacks[status] = this.#callbacks[status] || [];
this.#callbacks[status].push(callback);
}
record(success) {
const now = Date.now();
this.#events.push({ time: now, success });
this.#prune(now);
const newStatus = this.#computeStatus();
if (newStatus !== this.lastStatus) {
this.lastStatus = newStatus;
for (const cb of this.#callbacks[newStatus] || []) {
cb(this.report());
}
}
}
#prune(now) {
const cutoff = now - this.#config.windowMs;
while (this.#events.length && this.#events[0].time < cutoff) {
this.#events.shift();
}
}
#computeStatus() {
const frac = this.remainingFraction;
if (frac <= 0) return "exhausted";
if (frac < this.#config.criticalThreshold) return "critical";
if (frac < this.#config.warningThreshold) return "warning";
return "healthy";
}
get total() { return this.#events.length; }
get successes() { return this.#events.filter((e) => e.success).length; }
get failures() { return this.#events.filter((e) => !e.success).length; }
get currentRate() { return this.total ? this.successes / this.total : 1; }
get budgetTotal() {
return this.total * (1 - this.#config.targetRate);
}
get budgetRemaining() {
return Math.max(0, this.budgetTotal - this.failures);
}
get remainingFraction() {
const bt = this.budgetTotal;
if (bt <= 0) return this.failures === 0 ? 1 : 0;
return Math.max(0, this.budgetRemaining / bt);
}
get burnRate() {
const expected = this.total * (1 - this.#config.targetRate);
return expected > 0 ? this.failures / expected : 0;
}
report() {
return {
status: this.lastStatus,
currentRate: Math.round(this.currentRate * 10000) / 10000,
total: this.total,
failures: this.failures,
budgetRemainingPct: Math.round(this.remainingFraction * 1000) / 10,
burnRate: Math.round(this.burnRate * 100) / 100,
};
}
}
// Usage
const budget = new ErrorBudgetTracker({ targetRate: 0.95, windowMs: 3600_000 });
budget.on("warning", (r) => console.log(`[WARN] ${r.budgetRemainingPct}% budget left`));
budget.on("exhausted", (r) => console.log("[ALERT] Budget exhausted!"));
// Record results from your solver
budget.record(true); // success
budget.record(false); // failure
console.log(budget.report());
Un escenario real: una agencia que factura en USD
Una agencia en México o Argentina monitoriza precios en marketplaces y paga sus herramientas en dólares. Con un plan por threads de CaptchaAI —BASIC ($15/mes, 5 threads) o ADVANCE ($90/mes, 50 threads)— el coste mensual es predecible, porque cada thread incluye resoluciones ilimitadas. Cuando la tasa de éxito de reCAPTCHA v2 se desploma, el saldo se agota y el rastreador pausa esa cola en vez de ocupar threads contra un objetivo roto, dejándolos libres para las colas sanas (Turnstile, GeeTest v3).
Errores frecuentes y cómo resolverlos
Causas habituales de un saldo que se comporta raro:
| Problema | Causa | Solución |
|---|---|---|
| El presupuesto se agota demasiado rápido | SLO demasiado exigente para las condiciones reales | Fija un SLO realista basado en datos históricos |
| El presupuesto nunca se consume | SLO demasiado laxo | Ajusta el SLO para forzar mejoras de fiabilidad |
| El estado oscila entre valores | Ventana demasiado corta | Usa una ventana más larga (24 h en vez de 1 h) |
| Tasa de consumo engañosa con poco volumen | Pocos eventos distorsionan el cálculo | Exige un mínimo de eventos antes de calcular |
| La memoria del rastreador crece sin parar | Eventos sin podar | Comprueba que _prune se ejecute en cada record() |
Preguntas frecuentes
Dudas frecuentes al llevarlo a producción:
¿Conviene una ventana de 1 hora o de 24 horas?
Depende del uso. Una ventana corta (1 h) reacciona rápido, pero oscila mucho con poco volumen; una de 24 horas o 7 días es más estable de cara al SLO. Un patrón habitual es llevar ambas: la corta para alertas y la larga para los despliegues.
¿El presupuesto de errores encaja con la facturación por thread de CaptchaAI?
Sí. CaptchaAI cobra por thread concurrente con resoluciones ilimitadas, así que el gasto no crece con cada reintento. El presupuesto, en cambio, te dice cuándo dejar de gastar esos threads contra un objetivo que falla.
¿Debo llevar un presupuesto por cada tipo de CAPTCHA?
En general, sí. reCAPTCHA v2 puede tener un SLO del 93 % y Turnstile uno del 97 %; agregarlos en un único saldo esconde los problemas de cada tipo. Lleva un rastreador por tipo; súmalos solo en el panel, no en la lógica que decide cuándo frenar.
Próximos pasos
Mide la fiabilidad de tu resolución de CAPTCHA con cifras, no con intuiciones: obtén tu clave API de CaptchaAI e integra el rastreador en tu pipeline. Para profundizar:
- Seguimiento de resoluciones de CAPTCHA con DynamoDB sin servidor
- Patrón circuit breaker para llamadas a la API de CAPTCHA
- Endpoints de health check para workers de CAPTCHA
- Monitorización de tasas de resolución con Prometheus y Grafana