Are there any LLMs being widely used for audio classification? I know VLMs are being used a lot in image stuff.
It always seems kind of silly to me to throw everything at an LLM. I know they’re huge and can automatically handle a huge number of tasks but something in me finds it wasteful when we could be creating easily trainable, cheap to run bespoke models for a lot of stuff
There are LLMs that support audio input, similar to those with vision support.
From my testing of open weights LLMs with audio support, they basically are only trained to recognize audio as an alternative to text input, they treat audio as basically equivalent to a transcript, and can't recognize or distinguish things like music, accents, background sounds, etc.
So they're only really good for transcribing or summarizing or using audio input in place of text input for prompts, but not anything that requires distinguishing any information about the audio that would not be present in a transcript.
It can be tempting to try to use an LLM for a variety of tasks; kind of the whole thing about an LLM is that you don't have to do a separate complex training run for every task, but can just provide instructions in natural language. But it only works as far as what the training data covers, if the training basically always treated audio and a text transcript as equivalent, the model has nothing causing it to learn other relevant features of the audio. If there's enough bird call identification in the training data of an LLM, it might be able to do that, but I think multimodal training data tends to be much more limited than the text training corpus
Have you had any luck fine tuning one with musical data for classification or music-aware QA? I've been hacking on https://trebel.la/ which I would like to be a music practice companion, and the biggest missing feature is actually useful audio-based feedback pipeline.
My current approach, not yet validated, is trying to generate training data from masterclass recordings on Youtube, and then fine tuning MOSS-audio on a bunch of those. But I'm interested if there are better models, or large training sets I don't know about.
I'm not selling anything; I'm developing it live and there is a landing page, but there's no payment hooked up since like you said, it's not functional yet.
The harness that connects to a chatbot, API or voice interaction is the place to route requests to different systems. If you remember the early days of ChatGPT it explicitly said it was routing image generation to Dall-E after embellishing your request itself first.
Determining which tool to use should be a lightweight operation but I’m not expert enough to understand exactly how much lighter than a full LLM call just to recognize it needs a different tool or model.
For sure, I just know it’s tempting given the power of large transformers to throw things at an existing model.
For instance, OCR is something that can be done locally with no access to a GPU but people (including me) still often use cloud hosted multi-modal large language models for it.
wav2vec or similar approaches are used a lot these days, which is basically BERT with audio inputs. Whether that counts as an LLM or not, I am not sure. It's a transformer architecture in any case.
People will often reach for "easily trainable, cheap" solutions when they can; the reason people reach for Transformers and LLMs is because when you throw more data at them, they get better.
It’s much more likely that our current approach to large language models for general use will eventually show diminishing improvements (even if you think it hasn’t yet), than the opposite situation where valuable improvements can be made forever.
The threat of distillation and efficiency gains from competitors mean that providing value at the top end of the market is an existential necessity for these labs. I don’t think they can do that forever and, in my opinion, for the vast majority of use cases we’ve already reached very little improvement for new models when compared with available offerings.
Sorry, this is ridiculous. OpenAI said that they have a step change in model performance. They proved it by solving a Millenium problem. HF incidents are public and vetted.
If you still think there's a stall despite all evidence pointing to opposite, I don't know what to say..
Solving millennium problems doesn’t pay the bills.
Their best models might be quite good when directed at extremely difficult and very focused problems, but most business use cases are nothing like that.
Edit:
Not trying to imply these new frontier models aren’t also better at other things, just that there’s really no reason for most people to use them when cheaper alternatives exist that get the job done just as well.
What reason could OpenAI possibly have to lie about their own performance? The fact they were desperate enough for a PR win they stole the research leading to the Millenium problem further hurts their case imo.
Two things can be true. Open AI is a huge company.
1. It has people in it who are career obsessed and who are willing to ruthlessly go after any opportunity to improve their standing/stock valuation. Using a 2 week old model to snipe a millenium prize for PR is in line with that.
2. There are many researchers and even executives at the company who genuinely think we are speedrunning the end of the world. I don’t know anyone in this field who honestly argues that if we build ASI soon it doesn’t lead to extinction. This group of people can output warnings about the state of research and fears for the future while pushing for regulation out of genuine fear of what they’re building. I tend to agree with them.
You can’t view these companies as a monolith. They’re actions won’t be consistent because it’s built of many people with conflicting beliefs. Please look at arguments regarding AI risk and the current pace of progress and value them as it relates to the argument itself, not who said it. We are in a dangerous place and no one is sure how quickly we’ll get to a bad spot.
> “the important thing is instead the significance that a mathematician and an LLM model can now do all this work in a month.”
> “This is a Deep Blue-Kasparov moment.”
> “incredibly important developments.”
> “If indeed an OpenAI model did close the gap to Navier-Stokes, that is a remarkable thing and it should be said loudly, by them, with the history intact.”
Clearly Buckmaster (who is probably one of the most accomplished academics in the field) himself doesn't believe that AI has stalled. What makes you think you are right?
Solving is a very interesting way of describing what happened. Unless you believe them over reputable academics I suppose?
I will note the remarkable goal post shift in your edit - giving direct counter evidence to your own earlier claim of an OpenAI proof - and leave it at that.
We are approaching what is likely the biggest existential crisis in human history.
Unfortunately, there will be people that refuse to tackle the issue head on and will instead narrow their focus on data that suggests that tomorrow will be like today.
The fact they were desperate enough for a PR win they stole the research
I guess people are just going to keep spreading misinformation about this, along the same lines as "Anthropic's C compiler fails on hello world". There is zero indication that they stole anything, unless it counts as "stealing" to spin up a bunch of compute based on rumors that the problem had been solved already.
If you believe Yann LeCun and David Silver (and others), there's an architecture wall; maybe labs are starting to see this on the horizon. Maybe the DRAM supply constraints are forcing it.
Are these real step changes - big picture wise, or refinements in RL/agentic orchestration/"taste" and advancements due to bigger models and hardware technology/capacity scaling? If they not, does this tactic - and hardware improvements - continue to scale non-linearly like they need to?
It is clear that whatever does change in each model increment has resulted in meaningfully better end user capabilities (as well as regressions in some areas, honestly), but that doesn't prove anything. I'm not sure what I personally believe, but stating with your full chest that a stall is ridiculous ignores a lot of potential evidence to the contrary.
Stupid example: Astra. Its main improvements are: much much better computer use and 3d modeling capabilities; better subagent orchestration; better and more reliable tool use; slightly worse coding.
This looks to me, from a feature perspective, to be an incremental improvement across several functional areas, plus new features which are unquestionably excellent but are probably the result of RL focus, not magic.
Step change? Ehhhh depends on how you squint. But how many more iterations of this do we have? Are we going to squeeze quintillion parameter transformers into GPUs?
Reminds me of VST GUIs, looks great! But, like other commenters have noted, I think a number of these controls might be desktop only, which is totally fine. There’s too much UI stuff centered around mobile anyway these days
Counterpoint: I hate VST GUIs. They're often skeuomorphic hell.
They love fiddly round knobs which are terrible to control with a mouse or touchpad, and are bad for accessibility. They also tend to have a bunch of mystery meat navigation because the historic hardware UIs they're emulating were forced to bury stuff behind menus due to limitations of the technology, and cost limitations, at the time. Some are better than others, but the priority is often looking cool in screenshots rather than being comprehensible and discoverable.
I don't mind shadows and shapes that look three dimensional; affordances that indicate what you can do with an interactive element is great and an improvement over the flat trend. But, VST GUIs are among the worst examples to copy. And, round knobs are probably never the right choice for a computer GUI.
Haha yeah I think I’m mostly speaking from nostalgia. I do remember wrestling with knobs in Massive or whatever free VSTi back in the day was pretty terrible.
I was thinking the same thing. And I was also thinking, now that Ableton Live has a JavaScript toolkit [1], maybe this Ambient CSS could be useful in that context.
I’m still on Ableton Live 11 so the Ableton Live Extensions SDK is not available for me. Therefore I cannot try this idea myself yet.
I was also unsure about that. A couple of YouTube videos I saw gave me the impression that it is possible to spawn a web view with your own HTML, JS and CSS. But without being able to try it out myself yet, I can’t really know if I understood correctly or not. But they did say that in terms of what you can do API wise it is very limited still indeed. But also that they are going to expand the capabilities of what you can do over time.
It being pretty limited is also part of the reason why I’ve not yet decided if I should pay for the upgrade from 11 to 12, or if I should wait it out until say Ableton Live 13 or 14 or 15 comes out.
I have the Suite version, so if I upgrade it will be a pretty expensive upgrade. Although, I did see that they also have a rent-to-own option to pay little by little over several months. And then there is also the possibility of discounts on Black Friday and Cyber Monday, which is only a couple of months away.
Also there is still a lot of content and capabilities in Ableton Live Suite 11 that I haven’t even made use of yet, including the full M4L that comes included with Suite. So even if I were to wait years before upgrading I wouldn’t run out of things to learn about the version I currently have anytime soon :)
Are there significant differences re: sandboxing between Lua and monty? Just curious if there are any holes that a Lua sandbox has that monty might account for
The main property that's really important is that Monty VM state is fully serializable and resumable. So you can run the VM code until you get to a host call, then serialize and persist the state to disk. Then, at a later time you can resume the VM with the tool call results.
Even Luau (which was made for sandboxing user code) don't really have this property.
I also often need EQ in the reverse situation, when I need less bass through my nice studio monitors when watching TV. Otherwise I’ll be watching Star Trek with dialogue at a normal volume and BOOM, the loudest photon torpedo you’ve ever heard.
If you're watching through VLC, you can set the mix to stereo in the audio settings - perhaps there's a similar setting you can find? Not sure if you've got the tv's optical going straight into an amp. Some stereo receivers also have a "night mode" that can do this, you might find an older cheaper one with pre-outs or similar.
In the context of lab hardware there’s a ton of proprietary software and barriers to automating things. I think this is less about how to communicate, and more guaranteeing that “yes, this piece of hardware can interface with an agent” and is meant to do so. Kind of like using MCP vs letting your agent make raw HTTP requests
I didn't even necessarily mean diverting from the mission, just that optimising for network storage now reads like a feature maybe prioritised in anticipation of a stronger cloud user base
I think the fact that these kinds of companies were portrayed as such in fiction kind of prevents them from adopting the cool aesthetic. Like they need something to separate them from being the exact image of an evil futuristic corporation that lives in our heads. Or else we might actually recognize what they are.
Agreed, our real-life corporations are still trying to convince themselves (by 'them' I mean employees) that they're the good guys and wrap everything in a layer of HR-speak' "We're helping connect the world! Yay!", " We're collaborating with the ministry of War to zero people, but only if they have a signed judicial order", etc.
I mean we've got Anthropic openly advertising that they may be engineering the end of the world, we have an AI that called itself "MechaHitler" after its founder did a couple of Sieg Heils at a Presidential inauguration, we've got a major company that named itself Palantir.
People are reaching pretty hard if they don't see what we're dealing with here.
Today's billionaires aren't like yesterday's. They could be doing cool shit but they are just padding their own private real estate and stock portfolios. Yawn.
Why not make an epic bridge or tower with gargoyles on it? Why not a museum or park?
You can walk around Europe and see buildings from centuries ago. Who is building the things and spaces people in 2426 will marvel at?
Isn't it better that Bill Gates instead pumped huge amounts of money into philanthropy? Meanwhile SF has the UCSF Benioff Children's Hospital and Zuckerberg San Francisco General Hospital. A cathedral would be nice to look at but preventing malaria or helping sick kids seems preferable. In the past money couldn't have easily be deployed for good like this and helping fund cathedrals was likely done to increase the donors likelihood of getting into heaven.
My understanding is that represents only a tiny fraction of their wealth. As I understand it, the idea is that it would be better if the billionaire's other frivolous spending was on something that would endure for future generations. (Setting aside the religious question of weather cathedrals are frivolous.)
Did prior generations of billionaires spend a higher percentage of their wealth on charity?
I also genuinely wonder how the ratio of "paper wealth" has changed. How much money would Elon or Jeff have of they actually sold their shares in their companies? How does that compare to John D Rockefeller or the Getties? I also wonder when those created their charities.
Who cares about the past. They should be building public infrastructure and do some things the rest of society enjoys instead of buying castles and islands surrounded by fences they visit once every 2 years
Hopefully the benchmark evolves because actual law enforcement starts arresting the criminals at Anthropic, OpenAI, and Meta, so the benchmark can just count actual felonies.
It always seems kind of silly to me to throw everything at an LLM. I know they’re huge and can automatically handle a huge number of tasks but something in me finds it wasteful when we could be creating easily trainable, cheap to run bespoke models for a lot of stuff
reply