Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

Google OSS VRP Pause What It Means for STA Research

Sunday, October 04, 2026

Google's OSS VRP Pause: What It Means for STA Research and the Disclosure Ecosystem

On 1 October 2026, Google announced that it would stop accepting product vulnerability submissions to its Open Source Software Vulnerability Reward Program (OSS VRP). The reason: an overwhelming influx of AI-generated reports, including hallucinations, non-reproducible findings, and automated submissions that consumed engineering time without adding security value. The pause took effect the same day and is expected to last until at least Q1 2027.

This is not an isolated incident. Intel had already suspended its bug bounty program. Linux maintainers reported being flooded with AI-generated CVE reports, with some projects dropping support for older drivers as a direct consequence. The bug bounty ecosystem is buckling under the weight of automated noise.

But this post is not about those programs. It is about what the pause means for a specific, ongoing investigation, STA Research, and for the broader problem of how legitimate security findings survive in an environment saturated with AI-generated noise.


1. What exactly was suspended

Three distinctions matter:

Program Status Scope
OSS VRP: product vulnerabilities Suspended as of 1 Oct 2026 Code flaws, logic errors, design defects in Google's public repositories
OSS VRP: supply chain reports Active Dependencies, build pipelines, third-party components
Android VRP Active Android platform, Pixel firmware, Titan M security chip
Chrome VRP Active Chrome browser, Chromium
Cloud VRP Active Google Cloud products
Patch Rewards Program Active Proactive security improvements in open source

The suspension does not affect reports submitted before 1 October 2026, and it does not affect the Android VRP or Chrome VRP programs. Google has explicitly encouraged researchers to use Cloud VRP and Patch Rewards as alternative channels.


2. The case that matters: A-477279924

STA Research has an open case in the Google Android VRP: A-477279924, submitted on 20 January 2026. The case documents a class of defects in AOSP's handling of structured text across Binder boundaries, including unbounded TaskInfo serialization, missing size gates in Task.fillTaskInfo(), and the cascade of failures that follows when a single oversized task state is materialized repeatedly across multiple framework interfaces.

The case is currently blocked on an internal dependency (477593694), which according to my analysis corresponds to AOSP CL 3989977, a fix that exists internally at Google but has not been made public.

A-477279924 is not in the suspended program. It lives in Android VRP, not OSS VRP. But the pause matters because:

  1. Shared triage resources. The engineers and maintainers who process Android VRP reports are part of the same security organization that has been overwhelmed by OSS VRP submissions. The burnout described in Google's announcement is not confined to one program.
  2. Guilt by association. A 40-page technical report with multiple vectors, complete stack traces, and AOSP source verification is exactly the kind of submission that a triage engineer, saturated by AI-generated noise, might mistake for automation. The risk is not that the work is invalid. The risk is that the signal is lost in the noise.
  3. The fix is already internal. The fact that AOSP CL 3989977 exists means Google has already identified the defect. The question is whether the OSS VRP pause delays its public release, or whether it accelerates internal prioritization to close the loop without public disclosure.

3. Why this affects every legitimate researcher

The OSS VRP pause is a symptom of a larger problem: the cost of generating a report has collapsed, but the cost of validating one has not.

Before large language models, writing a vulnerability report required understanding the code, reproducing the behavior, and articulating the impact. The effort served as a natural filter. Today, a model can produce a 3,000-word report with plausible-sounding analysis and fabricated stack traces in seconds. The human reviewer still has to read it, verify it, and reject it, consuming the same time whether the report is real or hallucinated.

Google's response (pause the program, redesign the intake process, return in Q1 2027) is rational from a resource-management perspective. But it creates a perverse incentive: the noise wins by silencing the signal. Legitimate researchers whose reports are already in the queue, or who discover new vulnerabilities during the pause, have no clear path to disclosure through the affected program.

The alternatives Google suggests (Cloud VRP, Patch Rewards) are valid, but they do not cover the same scope. A finding in AOSP's frameworks/base that affects Android devices is not automatically a Cloud VRP issue. The Patch Rewards Program rewards proactive fixes, not vulnerability reports. The gap is real.


4. What this means for STA Research

STA has always been a long-term investigation. The initial report was submitted in January 2026. The finding has been extended through 8 reproduction sessions, 31 documented crash instances, and verification against AOSP source code. The work predates the OSS VRP pause by months.

Three practical consequences:

1. The Android VRP case remains active, but patience is required.
A-477279924 is not suspended. The internal dependency (477593694) is the bottleneck, not the program status. Follow-up should continue, but expectations should be calibrated to the broader organizational pressure.

2. Public documentation becomes more important, not less.
When official channels are congested, the blog and the GitHub repository become the primary record. The technical report is published. The PoC is published. The evidence is timestamped. The absence of a public patch does not mean the finding is invalid. It means the system is slow.

3. Alternative disclosure channels matter.
The CVE assignment process through INCIBE to MITRE is independent of Google's VRP infrastructure. The HackerOne dispute for Xiaomi (#4064049) follows a different path. These are not substitutes for Android VRP, but they are parallel tracks that do not depend on the paused program.


5. The bigger question: how does disclosure survive AI noise?

The OSS VRP pause is not the end of bug bounties. It is an inflection point. The industry is being forced to redesign intake, triage, and validation processes around a new reality: automated report generation is cheap, and human review is not.

Some possible directions:

  • Proof-of-work for reports. Require a reproducible PoC, a signed environment fingerprint, or a hash-verified test case before a report enters the queue.
  • Tiered intake. Separate "first-pass automated scan" reports from "human-verified, reproducible" reports, with different processing paths.
  • Reputation systems. Weight reports by the researcher's history of valid findings, with new submitters routed through stricter validation.
  • AI-assisted triage. Use models to pre-filter obvious hallucinations, but with the understanding that this creates new failure modes.

None of these are perfect. All of them shift the burden somewhere. But the current model, where a single invalid report costs the same to process as a valid one, is not sustainable.


6. What this means for an independent researcher

There is a version of this article that would end at section 5, with the systemic analysis complete and the policy recommendations neatly laid out. That version would be accurate. It would also be incomplete, because it would leave out the part that matters most to those of us who do this work without institutional backing.

The asymmetry was already there before AI. Researchers working at large security firms have salaries, legal teams, PR departments, and colleagues who can validate a finding before it goes out. An independent researcher has a laptop, a test device, and a blog. When a report is ignored, misclassified, or closed without explanation, the institutional researcher moves on to the next thing on their employer's roadmap. The independent researcher carries the cost personally: the hours spent, the reputational risk of publishing, the ambiguity of not knowing whether the work will ever be acknowledged.

AI noise did not create this asymmetry. It made it worse.

The verification burden falls hardest on those with the least capacity to bear it. When triage teams are overwhelmed, the natural response is to apply stricter filters. That is rational. But stricter filters favor submissions that come with institutional credibility markers: company letterheads, internal references, coordinated timelines. A careful independent report, correctly evidenced but unadorned, is more likely to be filtered out than a mediocre report from a known entity. The playing field tilts further, not less.

And then there is the money. Or rather, the absence of it. Bug bounty payouts are inconsistent, but they exist. CVE assignment does not pay anything. Blog readers do not fund research. This work is not a job. There is no salary, no retainer, no consulting fee waiting at the end of a successful reproduction. When the STA investigation was submitted to Google in January 2026, the acknowledgment came with a partial payment of $250. That amount does not begin to cover the hours invested: 8 reproduction sessions across 4 weeks, 31 documented crash instances, 8 bugreports collected by hand, hundreds of logcat lines read one at a time, AOSP source verified commit by commit against android.googlesource.com. If this were billed at any professional rate, the invoice would be five figures. It was not billed. It was never going to be.

This is not a complaint about a specific payout being too low. It is a description of how the economics actually work. The bug bounty programs exist to incentivize disclosure, and for some researchers at some companies, they do. But for the independent researcher working on a class of defects that do not fit neatly into a single CVE, the payout, when it comes, is a symbolic acknowledgment, not compensation. The reward is the recognition. If the recognition does not come, nothing does.

Reputation becomes the only currency, and it is hard to build from scratch. For someone with prior credits from Microsoft, Google, and Mozilla, a closed case is a setback but not a career risk. For someone starting today, without those credits, a single misclassification can be the end of the attempt. The OSS VRP pause raises the barrier to entry not by making the work harder, but by making the acknowledgment of the work less likely. That is a cost the industry is paying without fully pricing it in.

Public documentation is the only defense that does not depend on anyone's goodwill. This is why I keep writing the blog. This is why the GitHub repository exists. This is why every bugreport, every stack trace, every timestamp, every hash is preserved in the open. Not because publication is a threat to anyone, but because publication is the only record that survives a program's silence. When Google pauses OSS VRP, when Xiaomi closes a case as "not applicable" while confirming an internal ticket, when MITRE takes months to assign a CVE, the blog is what remains. It is not a substitute for official acknowledgment, but it is not dependent on it either.

What sustains independent research is not the payout. It is the craft. It is the satisfaction of finding something real, of documenting it correctly, of leaving a trail that another researcher can follow. It is the conversations with people who understand why a 124-byte transaction failing in the same binder session as a 1.1 MB transaction is interesting, when almost no one else does. It is the small community of researchers who recognize each other's work not because a program validated it, but because the work itself stands on its own.

I am not writing this to complain. I am writing it because the current conversation about AI noise in bug bounties is happening mostly among people whose institutional positions protect them from the consequences. The companies will adapt. The programs will redesign. The platforms will survive. The question that is not being asked loudly enough is: what happens to the researchers who cannot wait for the redesign?

Some will move on to other work. Some will publish less. Some will stop. The loss will not be visible in any quarterly report, because the value of a finding that was never made is not measurable. But it is real, and it is being priced into the ecosystem silently.

For my part, I intend to keep going. STA has too much unfinished business: the CVE assignment through INCIBE and MITRE, the HackerOne dispute with Xiaomi, the follow-up on A-477279924, the reproduction sessions that keep adding new evidence. None of that depends on Google's VRP status. All of it depends on doing the work and documenting it honestly. And none of it pays in any direct sense. That is not the point. The point is that the work is worth doing, and someone has to do it.

The pause is a setback. It is not the end of the investigation.


7. What I am doing now

While the OSS VRP pause runs its course:

  • The Android VRP case remains open. I will continue to follow up on A-477279924 and the internal dependency.
  • Public documentation continues. The blog and GitHub repository are the primary record of the investigation.
  • CVE assignment proceeds through INCIBE to MITRE. This path does not depend on Google's VRP status.
  • The HackerOne dispute for Xiaomi (#4064049) remains active. The contradiction between "not applicable" and the internal ticket XS-112757 is documented.
  • New reproduction sessions continue. The finding is not static; the evidence base grows.

8. A note on the irony

The STA research documents a class of defects that arise when structured input is allowed to propagate without size constraints across processing layers. The OSS VRP pause is a different manifestation of the same underlying pattern: an unconstrained input channel (automated report generation) overwhelming a downstream processing layer (human triage), producing a cascading failure (program suspension).

The difference is that in STA, the fix is technical: bound the field, gate the transaction, degrade gracefully. In the bug bounty ecosystem, the fix is procedural: redesign intake, separate signal from noise, restore the economic asymmetry between generating a report and validating one.

Both problems are solvable. Neither will be solved by pretending the input is not there.

El dilema de la divulgación coordinada

Monday, August 31, 2026

Cuando la responsabilidad es unilateral: el dilema de la divulgación coordinada

Una reflexión sobre el modelo actual de seguridad, sus asimetrías y sus consecuencias

Imagina la siguiente situación:

Has pasado semanas, quizás meses, investigando un comportamiento extraño en un sistema ampliamente utilizado. Has reproducido el problema en diferentes dispositivos. Has capturado logs, stacktraces, métricas de sistema. Has documentado cada paso con precisión. Has preparado un informe que cualquier ingeniero podría seguir para verificar el problema por sí mismo.

Envías el reporte al fabricante. Esperas. Recibes una respuesta automática. Semanas después, alguien te pide más información. La proporcionas. Vuelves a esperar.

Finalmente, recibes una respuesta:

“Hemos revisado tu informe y determinado que no cumple con los criterios para ser considerado un bug de seguridad.”

“Este problema pertenece a otro equipo.”

“Está fuera del alcance de nuestro programa de recompensas.”

“Por favor, utiliza el feedback in-product para reportarlo.”

El problema sigue existiendo. Los usuarios siguen expuestos. Pero la responsabilidad ha quedado diluida en un laberinto de equipos, programas y criterios.

Esta historia es más común de lo que muchos creen. Y revela una asimetría estructural en el modelo de divulgación coordinada que merece un análisis profundo.

Este artículo no trata sobre una vulnerabilidad concreta. Trata sobre el sistema que la gestiona o, más precisamente, sobre el sistema que a menudo no la gestiona.


1. El contrato implícito de la divulgación responsable

La divulgación responsable, también llamada coordinada, se ha establecido como el estándar ético en la seguridad informática. Su premisa es sencilla y, en apariencia, incuestionable:

“Si conocemos un problema que puede afectar a otros, debemos dar al fabricante la oportunidad de solucionarlo antes de hacerlo público.”

Esta lógica protege a los usuarios. Permite que las empresas corrijan vulnerabilidades sin exponer a sus clientes a ataques mientras el parche está en desarrollo. Es un modelo que, en teoría, beneficia a todas las partes.

En la práctica, el investigador acepta un conjunto de obligaciones que incluyen:

  • Reproducir el problema de forma fiable y documentada.
  • Proporcionar evidencia técnica suficiente (logs, trazas, código, pasos).
  • Evitar divulgar prematuramente para no poner en riesgo a los usuarios.
  • Informar al fabricante a través de los canales establecidos.
  • Facilitar la investigación con información adicional cuando se solicita.
  • Respetar los plazos de coordinación que la empresa propone.
  • Permitir que el proveedor prepare una solución antes de la publicación.
  • Documentar sus conclusiones de forma responsable y precisa.

Y, en muchos casos, el investigador hace todo esto sin ninguna garantía de reconocimiento, parche o recompensa. Lo hace porque cree en el modelo. Porque entiende que la seguridad es una responsabilidad compartida.

La lógica es impecable. Pero esa misma lógica debería funcionar en ambas direcciones.


2. El problema aparece cuando nadie es responsable y a la vez, lo son todas las partes inplicadas

Durante una investigación pueden aparecer problemas que atraviesan diferentes capas de un sistema. Un mismo comportamiento puede involucrar:

  • Aplicación (el software que el usuario ve)
  • Framework (la capa intermedia que soporta la aplicación)
  • Biblioteca nativa (código de bajo nivel, a menudo en C/C++)
  • Sistema operativo (el núcleo del sistema)
  • Fabricante / OEM (personalizaciones del sistema)

Y también:

  • Producto A (ej. Chrome, Firefox, Edge)
  • Producto B (ej. Android)
  • Componente compartido (ej. libminikin)
  • Infraestructura común (ej. Binder, IPC)
  • Servicio en la nube (ej. Llm's API)

Entonces aparece el fenómeno conocido por muchos investigadores:

“No es nuestro problema.”

Un equipo o vendor, puede indicar: “Esto es un problema de Android.”
Android puede responder: “No está dentro del alcance de nuestro programa.”
Otro equipo puede añadir: “Debe reportarse al producto correspondiente.”

Y el investigador vuelve al punto de partida.

El atacante no necesita saber qué equipo es responsable. El investigador tampoco debería tener que resolver el organigrama interno de una multinacional para encontrar al responsable. Si el problema atraviesa capas, el atacante ve un sistema. El investigador ve un sistema. La organización, sin embargo, puede verlo como tres equipos o más distintos.


3. El caso STA: una investigación transversal

Mi investigación sobre Structured Text Amplification (STA) comenzó en 2022, estudiando comportamientos relacionados con texto estructurado y agotamiento de recursos en Android. Lo que parecía un problema aislado en una biblioteca fue revelando un patrón más amplio.

Con el tiempo, aparecieron diferentes manifestaciones en distintos componentes:

Componente Síntoma Mecanismo
libminikin.so ANR, bloqueo del hilo principal Knuth-Plass O(n²)
Binder / SavedState TransactionTooLargeException, crash loops Serialización O(n²)
Llm's (modelo) Instruction Drift, generación de contenido sin contexto Atención O(n²)
Llm's (cliente) ANR, UI freeze libminikin O(n²)
Navegadores Bloqueo de renderizado Algoritmos de layout O(n²)

Lo interesante no era cada fallo individual, sino la posibilidad de que existiera un patrón común:

Entrada estructurada (texto repetitivo, baja entropía)
              ↓
    Transformación (tokenización, layout, serialización)
              ↓
    Amplificación del coste (algoritmo O(n²))
              ↓
    Agotamiento de recursos (CPU, memoria, tiempo)
              ↓
    Pérdida de disponibilidad (ANR, crash, DoS)

Este patrón aparecía en el cliente Android (la apps de Google, Mozilla, Meta, Microsoft, Xiaomi, entre otros). Aparecía en el sistema operativo (libminikin, Binder). Y, más tarde, apareció también en los Llm's, tanto en el modelo (pérdida de contexto) como en el cliente (ANR al renderizar respuestas largas o tareas simples como contar caracteres ).

El problema era real, reproducible y estaba documentado con stacktraces, métricas de sistema y pasos concretos. Pero al intentar reportarlo siguiendo los cauces establecidos, ocurrió lo que muchos investigadores han vivido:

  • VRP's: “Fuera de alcance.”
  • llm's VRP: “Bypass de guardrail de seguridad. Fuera de alcance.”
  • Feedback in-product: Canal adecuado, pero sin garantía de respuesta o mitigación.

El patrón STA existía. Las evidencias eran sólidas. Pero la responsabilidad quedaba diluida entre equipos, programas y criterios.


4. La anatomía de una derivación

Para entender el problema, es útil analizar qué ocurre cuando un reporte atraviesa el sistema de gestión de vulnerabilidades de una gran organización.

Fase 1: Recepción
El investigador envía un informe detallado. Recibe un acuse de recibo automático. El reporte entra en una cola de triaje.

Fase 2: Triaje inicial
Un revisor, a menudo con poco tiempo y muchos reportes, clasifica el problema. Si encaja en un patrón conocido, puede ser asignado a un equipo. Si no, puede ser rechazado por “falta de información” o “no reproducible”.

Fase 3: Análisis técnico
El equipo asignado analiza el problema. Si el equipo es el correcto, la investigación avanza. Si el problema cruza fronteras, aparece la pregunta: “¿Es realmente nuestro?”

Fase 4: Derivación
El problema se traslada a otro equipo. Ese equipo, a su vez, puede derivarlo a otro. Cada derivación reinicia parcialmente el proceso. Cada equipo aplica sus propios criterios.

Fase 5: Decisión final
En algún punto, el problema es clasificado como “fuera de alcance”, “no elegible para recompensa” o “no reproducible”. El investigador recibe una respuesta. El problema sigue existiendo.

Lo paradójico es que cada decisión individual puede ser razonable. Cada equipo puede tener argumentos válidos para no asumir la responsabilidad. Pero el resultado final es que el problema no se soluciona.

Y el investigador, que empezó con la intención de ayudar, se encuentra con un muro de silencio.


5. “Out of scope” no significa “el problema no existe”

Hay una confusión conceptual que conviene aclarar.

Un programa de recompensas puede establecer legítimamente qué tipos de problemas son elegibles para recompensa. Esa es una decisión de alcance. Es razonable que una empresa defina los límites de su programa.

Pero:

No elegible para recompensa ≠ inexistente.

Un problema puede quedar fuera de un VRP y seguir siendo:

  • Reproducible.
  • Técnicamente relevante.
  • Peligroso para determinados usuarios.
  • Digno de una mitigación.
  • Digno de una investigación interna.
  • Digno de ser documentado públicamente.

Esta distinción es fundamental. Un programa de recompensas puede rechazar un reporte por alcance, pero eso no significa que el equipo de producto deba ignorarlo.

El problema ocurre cuando “fuera de alcance” se convierte en un sinónimo de “no es responsabilidad nuestra” y cuando esa falta de responsabilidad impide que el problema se solucione.


6. El coste de la investigación independiente

Para entender la asimetría, hay que considerar los recursos de cada parte.

Una gran organización puede disponer de:

  • Equipos especializados en diferentes áreas.
  • Acceso al código fuente completo.
  • Infraestructura de reproducción a gran escala.
  • Telemetría para identificar la prevalencia del problema.
  • Ingenieros dedicados a tiempo completo.
  • Herramientas internas de análisis y depuración.
  • Capacidad para parchear millones de dispositivos en días o semanas.
  • Departamento legal para gestionar riesgos.
  • Presupuesto para recompensas y reconocimiento.

El investigador independiente, en cambio, puede disponer de:

  • Un ordenador (a menudo personal).
  • Un teléfono (a menudo personal).
  • Unos bugreport (obtenidos con esfuerzo).
  • Una conexión a Internet.
  • Y muchas horas de trabajo no remunerado.

En mi caso, buena parte de esta investigación se ha realizado desde un entorno doméstico. No hay un laboratorio detrás, ni un departamento legal, ni un equipo de ingeniería esperando para validar cada hipótesis. La validación de las evidencias recae enteramente en el investigador.

Y, sin embargo, el investigador debe proporcionar evidencia suficientemente sólida para que una organización pueda tomar una decisión. La exigencia es legítima. La reciprocidad debería serlo también.


7. Cuando la evidencia contradice la respuesta inicial

Una de las situaciones más reveladoras ocurre cuando la primera conclusión de la organización es:

“No reproducible.”

Pero posteriormente aparecen:

  • Nuevos dispositivos donde el problema se manifiesta.
  • Nuevos dumps con stacktraces adicionales.
  • Nuevos ANR traces en el mismo componente.
  • Nuevas aplicaciones afectadas por el mismo patrón.
  • Nuevas reproducciones que confirman la hipótesis.
  • Evidencia del mismo componente en diferentes contextos.
  • Comportamiento consistente entre productos.

Entonces la pregunta ya no debería ser:

“¿Por qué el investigador insiste?”

La pregunta debería ser:

“¿Qué hemos aprendido desde la primera evaluación?”

La seguridad no debería funcionar como un juicio que termina con la primera decisión. Debería funcionar como un proceso iterativo:

Hipótesis inicial
        ↓
Evidencia recopilada
        ↓
Reproducción en condiciones controladas
        ↓
Análisis técnico
        ↓
Nueva evidencia (más dispositivos, más contextos)
        ↓
Reevaluación de la hipótesis
        ↓
Actualización de la decisión

Este ciclo es común en la investigación científica. En la seguridad, sin embargo, tiende a ser lineal: una decisión inicial, sin espacio para la reevaluación.


8. La paradoja de la coordinación

Cuando una organización solicita coordinación, el mensaje es claro:

“Danos tiempo para investigar y solucionar el problema.”

El investigador acepta. Pero la coordinación implica una segunda obligación: utilizar ese tiempo de forma efectiva.

La coordinación no debería significar:

Investigador
    ↓
Reporte (con evidencia)
    ↓
Espera (semanas o meses)
    ↓
"No reproducible"
    ↓
Investigador aporta más evidencia
    ↓
Espera
    ↓
"Out of scope"
    ↓
Investigador apela
    ↓
Espera
    ↓
"Pertenece a otro equipo"
    ↓
Investigador reporta al otro equipo
    ↓
El ciclo se reinicia

Eso no es coordinación. Es derivación de responsabilidad. Es un laberinto donde el investigador es el único que recorre todas las salas, mientras la organización mantiene sus puertas cerradas.


9. Una contradicción evidente

Al investigador se le dice:

“No publiques todavía.”

Perfecto. Es razonable.

Pero si después de meses o años la respuesta continúa siendo:

“No es nuestro problema.”

¿Durante cuánto tiempo debe permanecer el investigador en silencio?

  • ¿Quién protege al usuario durante ese periodo?
  • ¿Quién asume el riesgo de que el problema sea explotado?
  • ¿Quién decide que el problema merece atención?
  • ¿Quién determina si el problema es “suficientemente grave”?
  • ¿Dónde termina la responsabilidad del investigador y empieza la responsabilidad del fabricante?

El modelo actual responde a estas preguntas de forma implícita:

“El investigador es responsable de no divulgar. El fabricante es responsable de decidir si el problema existe.”

Pero la decisión de “si el problema existe” no debería ser una decisión unilateral, especialmente cuando el investigador ha aportado evidencia sólida y reproducible.


10. La responsabilidad no puede viajar solo en una dirección

El modelo actual puede resumirse así:

Investigador Organización
Reproducir el problema Investigar técnicamente
Documentar con evidencias Validar la información
Reportar a través de los canales Responder en tiempo razonable
Coordinar la divulgación Coordinar la corrección
Esperar el tiempo necesario Actuar sobre el problema
No divulgar prematuramente Mitigar el riesgo
Facilitar información adicional Asumir responsabilidad

El problema aparece cuando la segunda columna se convierte en:

“No corresponde a nuestro programa.”

“No es elegible para recompensa.”

“Pertenece a otro equipo.”

Entonces la primera columna sigue teniendo todas las obligaciones, mientras que la segunda conserva únicamente la posibilidad de rechazar el caso.

Eso es una asimetría estructural. No es un fallo de una empresa concreta. Es un fallo del modelo.


11. La recompensa tampoco debería ser el centro

Hay una cuestión especialmente importante que suele pasarse por alto.

La investigación de seguridad no debería reducirse a:

Bug → CVE → recompensa

Hay investigadores que buscan dinero. Otros buscan reconocimiento. Otros simplemente quieren que el problema se arregle. Algunos investigan porque quieren comprender cómo funcionan los sistemas y compartir ese conocimiento.

Por eso una respuesta como:

“No es elegible para recompensa”

no debería cerrar necesariamente la conversación técnica.

Podría existir otra respuesta:

“No podemos recompensarlo según las reglas del programa, pero hemos identificado el problema y vamos a mitigarlo.”

“Hemos derivado el problema al equipo de producto para que lo evalúe en futuras versiones.”

Esa sería una respuesta mucho más saludable para el ecosistema.


12. El silencio como estrategia

Hay una realidad incómoda que pocos investigadores mencionan abiertamente.

En algunos casos, el silencio —o la derivación, no es un fallo del sistema, sino una estrategia deliberada.

Si un problema no se clasifica como vulnerabilidad, no hay obligación de parchearlo.
Si el problema se deriva a otro equipo, la responsabilidad queda en suspenso.
Si el investigador se cansa y desiste, el problema desaparece del radar.

Esta estrategia no requiere mala fe. Puede ser simplemente el resultado de equipos que trabajan bajo presión, con recursos limitados, y que priorizan los problemas que encajan en sus métricas.

Pero el efecto es el mismo: el problema no se soluciona.


13. El investigador independiente no tiene voz en la decisión

Una de las asimetrías más profundas es la siguiente:

El investigador aporta el descubrimiento. Aporta la evidencia. Aporta el tiempo. Aporta la paciencia. Aporta la buena fe.

Pero no tiene voz en la decisión final.

  • No decide si el problema es “suficientemente grave”.
  • No decide si merece un parche.
  • No decide cuándo se solucionará.
  • No decide si se reconocerá su trabajo.
  • No decide si se comunicará públicamente.

La organización tiene todas esas decisiones. El investigador tiene solo la decisión de publicar o no publicar.

Y esa decisión, publicar, está cargada de riesgos: legales, reputacionales, y de relación con futuros reportes.


14. Divulgación coordinada no es silencio coordinado

Existe una diferencia esencial entre ambas cosas:

Divulgación coordinada:

“Tenemos un problema. Trabajemos juntos para entenderlo, mitigarlo y comunicarlo de forma responsable.”

Silencio coordinado:

“El problema está reportado, pero nadie quiere asumir la responsabilidad. El investigador espera. El problema sigue existiendo.”

La primera protege a los usuarios. La segunda protege principalmente al proceso.

La primera es colaboración. La segunda es inacción.

Y la seguridad debería estar diseñada para proteger a los usuarios, no los procesos internos.


15. El caso STA como ejemplo de un problema más amplio

STA no es una excepción. Es un ejemplo de lo que ocurre cuando un comportamiento atraviesa diferentes capas de un ecosistema y la responsabilidad queda fragmentada.

En mi investigación, el mismo patrón apareció en:

  • Android (libminikin, Binder, SavedState).
  • Llm's (modelo y cliente).
  • Aplicaciones de terceros (WhatsApp, navegadores).
  • Componentes compartidos (StaticLayout, LineBreaker).

Cada uno de estos dominios tiene sus propios equipos, sus propios programas de recompensas, sus propios criterios y sus propias prioridades.

Pero el patrón subyacente es el mismo. Es la misma entrada estructurada, la misma amplificación de coste, el mismo agotamiento de recursos.

Sin embargo, cuando intenté reportarlo de forma transversal, me encontré con que:

  • VRP's lo consideraron “fuera de alcance”.
  • llm's VRP lo consideraron “safety guardrail bypass”.
  • El feedback in-product no garantiza respuesta ni mitigación.
  • El problema sigue existiendo.

STA no es un problema de un equipo. Es un problema de arquitectura. Y los problemas de arquitectura no se solucionan derivando responsabilidades.


16. Lo que debería cambiar

Para que la divulgación coordinada funcione de forma efectiva, se necesitan algunos cambios en el modelo actual:

a. Puntos de entrada transversales
Las grandes organizaciones deberían tener puntos de entrada para problemas que cruzan equipos. Un equipo central de triaje que pueda evaluar un problema técnico sin necesidad de que el investigador conozca el organigrama interno.

b. Distinción clara entre “alcance” y “existencia”
Que un problema no sea elegible para recompensa no debería impedir que el equipo de producto lo evalúe y, si es necesario, lo mitigue.

c. Procesos de reevaluación
Si el investigador aporta evidencia adicional que contradice una decisión inicial, debería existir un proceso para reabrir la investigación sin necesidad de reiniciar todo el ciclo.

d. Comunicación transparente
Si el problema se deriva a otro equipo, el investigador debería ser informado de forma clara, con un punto de contacto o un identificador de seguimiento.

e. Reconocimiento sin recompensa
Si el problema no cumple los criterios de recompensa, pero es técnicamente relevante, la organización debería poder ofrecer un reconocimiento simbólico (mención en los agradecimientos, nota en las release notes, etc.).


17. Una pregunta incómoda (y su respuesta)

Después de años investigando vulnerabilidades, con mas de 400 descubiertas y documentadas y mas de 80 CVE, observando este patrón, hay una pregunta que considero inevitable:

¿Qué debe hacer un investigador cuando ha cumplido con todas las reglas de la divulgación responsable, pero ninguna organización acepta la responsabilidad de solucionar el problema?

No tengo una respuesta sencilla. Pero sí tengo una conclusión:

La responsabilidad no puede exigirse unilateralmente.

Si se espera que el investigador actúe responsablemente para proteger a los usuarios, las organizaciones deben hacer lo mismo. La seguridad no es un juego de trileros donde la responsabilidad se pasa de una mano a otra hasta que el investigador se cansa.

El investigador debe asumir su parte. Pero la organización también.


18. El objetivo final

No se trata de ganar una discusión. No se trata de conseguir una recompensa. No se trata de demostrar que una empresa se equivocó.

Se trata de algo mucho más sencillo:

Que el problema deje de existir.
  • Si una vulnerabilidad puede solucionarse, solucionémosla.
  • Si no es vulnerable, demostremos por qué.
  • Si está fuera del alcance de un programa, derivémosla al equipo adecuado.
  • Si el impacto no alcanza el umbral de una recompensa, eso no impide investigarla.

Pero no deberíamos permitir que el último paso sea:

“Este problema pertenece a otro.”

Porque entonces el problema sigue perteneciendo a todos. Y, al final, a nadie.


19. Una llamada a la responsabilidad compartida

La divulgación responsable nació como un pacto de confianza entre investigadores y fabricantes. Ese pacto sigue siendo necesario. Pero la confianza funciona en ambas direcciones.

El investigador debe asumir responsabilidad por lo que descubre.

  • Investigar con rigor.
  • Documentar con precisión.
  • Reportar con buena fe.
  • Coordinar con paciencia.
  • Divulgar con responsabilidad.

Las empresas deben asumir responsabilidad por lo que construyen.

  • Responder con seriedad.
  • Investigar cuando exista evidencia suficiente.
  • Distinguir entre “fuera de alcance” y “no existe”.
  • Evitar derivaciones infinitas.
  • Proporcionar puntos de contacto adecuados.
  • Informar cuando la investigación continúa.
  • Mitigar cuando sea necesario.

Cuando un investigador entrega evidencia reproducible, concede tiempo y respeta los mecanismos de coordinación, la respuesta no debería ser una cadena infinita de derivaciones.

Debería existir una puerta de entrada.

Alguien que diga:

“Entendido. Nosotros nos encargamos de averiguar quién debe solucionarlo.”

Porque esa es precisamente la diferencia entre gestionar un reporte y gestionar un riesgo de seguridad.


Conclusión: la seguridad no es un juego de trileros

La seguridad informática es un campo que se basa en la confianza. Confiamos en que los fabricantes corrigen los problemas que les reportamos. Confiamos en que los investigadores no explotan las vulnerabilidades antes de que se solucionen.

Pero la confianza no es un recurso infinito. Se agota cuando una de las partes no cumple su parte.

La divulgación coordinada no debería ser una excusa para que las empresas trasladen todo el riesgo al investigador. No debería ser un mecanismo para silenciar problemas incómodos. No debería ser un laberinto del que el investigador no pueda salir.

Debería ser un proceso colaborativo donde ambas partes asumen sus responsabilidades para proteger a los usuarios.

El investigador descubre. El fabricante corrige.

Y el usuario, al final, está protegido.

Ese es el objetivo. No deberíamos perderlo de vista.


Como final del artículo diré: si clicar en un enlace causa el crash wn una aplicación y esta aplicación hace caer SystemUI y a su vez causa un loop de reinicios y se la interfaz y obliga al sistema a borrar sus propios datos de estado etc y de la que un usuario normal no sabe recuperarse, no es un problema de seguridad entonces que es?

Manuel García Peña (Lostmon) — Independent Security Researcher
Agosto de 2026

STA - Structured Text Amplification In Llm's

Sunday, August 23, 2026

🧩 Structured Text Amplification (STA)

A Systemic Vulnerability in LLMs. Documented from the Couch
📅 August 23, 2026 👤 Lostmon 🏷️ Research / Vulnerability / Tokenization

Structured Text Amplification (STA) is a phenomenon where a finite-length input sequence, composed of non-semantic characters and lacking structural delimiters, causes a non-linear growth in computational cost in generative AI systems.

We tested 7 different systems (DeepSeek, Grok, Gemini, Copilot, Leo, Qwen VL, and others) and all are vulnerable, though with different symptoms: reasoning loops, 19-minute thinking times, parsing errors, interface amplification, and more.

🧠 The Key: STA is not a flaw in a specific model, but a structural problem in the design of AI systems — affecting tokenization, the ingestion interface, the parser, and the reasoning mode.

⚙️ STA Pattern Used

The base pattern is a repetition of special characters without separators or semantic meaning

We tested lengths of 1,600, 10,000, 61,560, and 65,560 characters, always with UTF-8 encoding.

📊 Results by System

SystemInputMain SymptomAmplification
DeepSeek (V4-Flash)1,600 charsReasoning loop, long responses~4.4x
Grok (xAI)1,600 chars"Think" mode activated for 19 min without response—
Gemini (Google)61,560 charsLong structured response, no useful data—
Copilot (GitHub)61,560 charsInflated count (1,002,682) + fragmentation16.28x
Leo (Mistral)65,560 charsSyntax error: "Unterminated string"—
Qwen VL 30B65,560 charsSyntax error: "Unterminated string"—

🧬 Layer-by-Layer Analysis

LayerVulnerabilityAffected Systems
Input ParserDoes not escape special characters → syntax errorLeo, Qwen VL
Ingestion InterfaceConverts long text into truncated document with repetitionsCopilot
Tokenizer (BPE)Fragments special characters as individual tokensDeepSeek, Grok, Gemini
Inference EngineEnters loop without semantic structureDeepSeek, Grok
Reasoning ModeSTA prevents convergence → prolonged blockGrok

🎯 Attack Vectors Identified

  • Direct Vector: Sending the STA pattern as a message to the model (1,600 characters).
  • File Vector: Uploading the pattern in a file (61,560 characters).
  • Interface Vector: Pasting the pattern into an interface that converts it to a document (Copilot).
  • Multimodal Vector: Including the pattern in an image/text context (Qwen VL).
  • Clipboard Vector: Fragmentation and contamination of the clipboard (Copilot).

🔍 The Copilot Case: Interface Amplification

Copilot does not amplify STA by itself; rather, the interface converts the long text into an internal document (<AttachedDocument>), truncates it, and fills it with repeated blocks. The model receives that amplified document and processes it as if it were real.

Actual input: 61,560 characters Internal document: ~1,002,682 characters Amplification factor: 16.28x

This is especially serious because the user has no control over this process, and the attack can escalate without the model or the user detecting it.

💰 Estimated Economic Impact

ModelInput (chars)Approx. Cost per Attack
DeepSeek1,600$0.0014
OpenAI GPT-4 (reference)1,600$0.06
Copilot (with amplification)61,560 → 1MNot quantified, but high

If the attack is automated (10 requests/second), costs can quickly escalate to hundreds of dollars per hour.

🛡️ Technical Recommendations

  • Parser: Automatically escape non-ASCII characters and validate string termination.
  • Ingestion Interface: Do not convert long text into internal documents if it is not an explicitly uploaded file. If converted, do not truncate with repetitions.
  • Tokenizer: Add subword merging rules for common combinations of special characters and limit the number of tokens per input.
  • Inference Engine: Implement timeouts in reasoning mode and detect low-entropy patterns to respond with an error without spending tokens on inference.
  • User: Do not paste long strings of special characters into AI interfaces; use preprocessing tools that clean non-semantic characters.

📌 Conclusions

  • STA is a real and documented phenomenon affecting generative AI systems across multiple layers.
  • All tested models are vulnerable, though with different symptoms.
  • The interface layer can amplify the attack (Copilot: 16.6x).
  • STA is not a flaw in a specific model, but a structural problem in the design of AI systems.
  • Urgent action is recommended to mitigate this attack vector.

📚 References & Further Reading


🛋️ Research led from the couch with ingenuity, patience, and insatiable curiosity.
Lostmon
🧩 STA — Structured Text Amplification  ·  Version 2.0 (Technical)  ·  Published under CC BY-NC 4.0
```
 

Browse

About:Me

My blog:http://lostmon.blogspot.com
Mail:Lostmon@gmail.com
Lostmon Google group
Lostmon@googlegroups.com

La curiosidad es lo que hace
mover la mente...

Friends