2 posts · 1 sub · RSS
hour hourdayweekmonthyearall
allr/LocalLLaMA
▲
6
 
1👁
r/LocalLLaMA · u/pmttyji · 42m ago
[Paper] DLoop: Looped Speculative Decoding
Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at this https URL.
The code is currently under internal review and will be released soon. Stay tuned!
▲
0
 
1👁
r/LocalLLaMA · u/MrCatberry · 49m ago
Reverse Engineering Web Application/Service

Hi Guys!

Is there a known harness/workflow that makes it easier to Reverse Enginerring a Web Application/Service thats behind a payed subscription?

My current problem is that I use a service that costs me quite a lot of money but still does not have all the features or tweaking options I need.

Now I'm asking myself, if it would be best to describe every feature myself or if there is a way to let a LLM "explore" the Web Application/Service by itself and write it's own notes what it needs to code.

Did somebody here do something familiar and has some tips?

Thanks in advance!