The InfoTech Podcast

Dr. Arshavir Blackwell - YourVoiceCraft

Jimmy Huber

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 31:04

Send us Fan Mail

In this episode, we explore the cutting-edge of AI with Dr. Arshavir Blackwell, a pioneer with 30 years of experience. We discuss local large language models, voice personalization, security, and the future of AI tools that can transform content creation and business workflows.


YourVoiceCraft


Dr. Arshavir Blackwell:

LinkedIn

Arshavir.com

SPEAKER_00

And we're back with another episode of the Info Tech Podcast. I'm your host, Jimmy Huber, as always. And it seems to be an AI trend these days. If you're a regular listener, um, you know that I've had multiple people and guests on talking about the advancements with artificial intelligence, the good, the bad, and the ugly. And today is no exception to that, but we have a twist. Um, my guest is uh a PhD uh in cognitive science. Um he's got 30 years of experience in AI. He's not only um, I think, forging ahead in the frontiers, but he's been in this industry much longer than anyone I've had on or have heard of, frankly. Um his name is Arshavir Blackwell, and I really appreciate you spending the time with us and welcome to the podcast today. Thank you. It's great to be here. Yeah, you bet. Um, we'll get to your app in a minute. Uh I think a lot of people are gonna enjoy hearing about that and even being able to try it out. But I want to try to kind of take you back in in time. I always want to give my audience the um the feel for who we're talking to here. So um if you wouldn't mind, Arshavir, take a couple of minutes and just tell us about your background. Obviously, you you've had a lot of years in this space, but um, how did you come to make this app and what kind of brought you here?

SPEAKER_01

Well, it's been a long trip. Um I was as a even as a kid, I was one of those kids that was always interested in computers and technology and science fiction. And so that led me to study cognitive science and computer science. And I got my doctoral degree, my graduate degree at UC San Diego, and I was really fortunate because I got to work with two of the real top people in the field at the time, Jeffrey Elman and Elizabeth Bates. And Elman was one of the people that developed some of the earliest neural network technologies that is, of course, quite primitive compared to what we're seeing today. But at the time, it was really quite revolutionary and it's definitely foundational. So a lot of the models that he created are the basis for some of the models that we see today actually employed in large language models. So after that, I worked for a series of mostly startups, and I my interest was in bringing neural networks to the masses, sort of. So I always saw that the technology had an enormous value and an enormous potential. And unfortunately, it was kind of cutting edge and there was a lot of controversy around it amongst some people. And so the problem was that uh, you know, to some extent, um, you had to kind of, with some people at least, fight to get your ideas out there. Um, so uh, you know, it was sometimes challenging because this was very new cutting-edge technology. But I think we did a lot of good work and I think we created a lot of really cool applications. So that's what led me basically to where I am now. I have a Substack where it's uh inside the blackbox.ai, and it's all about looking inside how these models work and understanding what's going on when you type in something and you get a response at the other end. So the surprise to a lot of people is that even the people that have built these things, they've built systems that are so incredibly complicated that they don't themselves understand really what's going on inside them. They know they put in input and they get out some reasonable output, but they don't really get, you know, if you change this one thing, what happens and why? We don't really have a good understanding of how they work yet. And so that's an area of great interest to me. The other area that I really find fascinating is uh the area of voice analysis and basically taking someone a set of writing samples from an individual and using that to create a sort of voice print so that the local language model can write in that person's voice. And I think we found, you know, a lot of people who've used large language models will say they're very powerful, but when they write, they all sound the same. And so when I try to write something, I end up having to rewrite it in my own voice, which kind of makes the whole point of using a large language model really ups up. Yeah, it defeats the purpose, exactly. So what this does, it's what's called a local large language model. And that's in contrast to the large language models most people are used to right now, are things like ChatGPT and Claude. And those are frontier language models. And what that means simply is that they're at the very edge, the cutting edge of technology in terms of what kinds of computations and algorithms and processes they can do. And they're also online. So they're never on your machine. What you're doing is basically connecting to a cloud service, typing something in, sending it out to that service, and then getting something back. A local LLM can be completely contained on your own machine if you want it to, or in a local server that's not Chat GPT or not Cloud or doesn't belong to Anthropic. And there are a lot of reasons why, you know, surprisingly, that's actually a better model. One of those is security. So a lot of people work with things like HIPAA compliance, legal compliance, and they don't really feel comfortable sending a lot of information that's supposedly secure information or that you know can't risk being hacked into out to anthropic or uh out to open AI. And they even say when you're using it, you know, this is not a confidential channel. Don't put passwords in here, uh, don't put anything in here that you know you would be afraid of having leak out accidentally. And of course, they also can use the data that you send them, the information that you send them as training data uh for their own models. So what you're working on can appear later on somewhere in their model. So that's not a great thing. Yeah. Whether you know it or not, they're not going to tell you, of course. Right. The other thing is, of course, those models really eat up money very quickly. So you know, if you're paying by the token and you use 100,000 tokens or something like that in the project, I mean it just it can become you can get a very large bill very quickly. Trevor Burrus, Jr. A lot of overhead. Sure. Trevor Burrus, Jr.: There's a lot of overhead. If you're using a local language model, especially if it's on your machine, it's free because you're just using your machine.

SPEAKER_00

So it's just like you're providing your own compute instead of farming it out to a data center or whatnot.

SPEAKER_01

Exactly. So you're not paying for that. And the other advantage of a local language model is it allows you to do this personalization, which you can't do with frontier models. So there's no way to change the underlying mechanism of the frontier model so that it sounds more like you. There's no knobs or buttons or anything like that to kind of tweak. With the local language model, you can put what's called an adapter on it. And the adapter learns your voice from what it is you've written, and it then starts to take on that exact voice when it writes new material. And so I've done, I've run some examples on really strong voices. Uh and if you look at, say I've trained it on Twain, I've trained it on Edgar Allan Poe, uh, I've trained it on Shakespeare, I've trained it on a lot of different very well-known authors. And if you just take, and you can actually see this on the website, if you just, if you type something in and you have it write as Twain, it sounds extremely different than if you have it write as Poe or Shakespeare or Hemingway. And the reas the reason this doesn't work on Frontier models is because the frontier model hasn't really been trained to absorb the actual cadence and the actual sort of the metrics that make up your voice. So what it does is it puts on a costume. It'll say, if you're doing Twain, for example, it'll say uh I I reckon and let's and things like that. You know, let's go, let's go yonder. But it's not going to actually clunky. Yes, it's very clunky. It's not going to take on your the style of Twain or Poe or of you. And that's one of the big advantages of using the local large language model is it permits you to do a lot of personalization in a really big way. And we're finding now just all sorts of new applications and uses for it once you've captured that individual's voice and their work product.

SPEAKER_00

Yeah. So let me let me back up just a little bit because you sure you said a ton of stuff there that I want to make sure our audience understands. So I agree that I think most people are used to Chat GPT, Gemini, you know, they may not know it's an LLM, but they know that, hey, this is the new Google and it seems to be a little bit better, right? Like just skimming the surface. Or you've got some professionals, young professionals, maybe or business leaders that are trying to augment it. But yeah, like you said, that it spends more time training or to make it even passable than the time it maybe saves you. Or the results just aren't good enough for prime time. So you got a lot of dabbling, a lot of testing, et cetera. Your your app here um specifically takes away the model that all those work in and says, hey, you can take the the value of a of an LLM, you can run it locally on what, your phone, your desktop computer, or whatnot. Yep. And you can actually spend a little time training it and giving it information that you've generated itself, and it's going to be what, 50% better, 20% better? I mean, we we talked numbers, I think, last time a little bit. What what are what can be the expectation here? Because I would think some people say, well, I'm doing that already with Claude or something. It knows me really well. You know, what do you say to that?

SPEAKER_01

This will give you a much better result. So in some of the tests that I've run, let's say that you give Claude a passage and ask it to uh write it in the style of, for example, Mark Twain. If you evaluate that, let's say based on some metrics, the different stylistic metrics, stylometric metrics that we use, um, that it gives you like a a score of 45 when a base would be a 20. What we find when we run it through the LLM that's the local model, when we actually train it, it's let's say 75 or 80. And there's no real mystery as to why that is. It's because you're actually changing the weights of the system so that all of the decisions that it's making have been changed at a very fundamental level. In the frontier models, the chat GPTs and so forth, you don't have that. What you're doing is you're giving it just a verbal command, right like Twain. And it's kind of putting a costume on what it would have done anyway. And that's the reason why there's such a big difference. So this is an interesting example of where frontier models are really good for a lot of things, but it's an interesting example of where local smaller language models are actually better. And the thing that I always like to sort of point out to people is you while they're very powerful, you don't always need a frontier model, a super large frontier model. So it's like you don't, if you have a destroyer or a speed boat, the speedboat will get you across the lake just as, you know, well as the destroyer. So the frontier model is very powerful and has a lot of things that the small model can't do. But in this one particular use case we're finding, it actually is not only as good, but it's actually quite a bit better.

SPEAKER_00

Yeah. I think you hit the nail on the head there. Most people, I don't think, realize what they're writing in at all if they even know they're on the water. Yeah to take your analogy further. So you're you're you're attacking the specific, probably overarching mainstream frustration with AI is that it doesn't sound like me. It's not actually helpful. I would never send information to this person as me because it it's not me, right? And people are already starting to pick up on that. So you're saying, hey, you can take it local, but you don't need a ton of resources because you're you're simplifying it. You're giving them the speed boat, right? Which is all they need essentially anyway. So that's I would think would probably be a lot cheaper then too, because most people are figuring out unless you get a paid version of chat or Gemini or whatever, they kind of suck, right? Like you have to invest at a certain point. And I think those price points might go up as the, at least at the enterprise level, the token usage is just nuts. Uh all the things I'm reading now are people are realizing, hey, it it actually is cheaper to hire humans at scale, right? So you're you're attacking and addressing multiple of these things.

SPEAKER_01

Absolutely. Absolutely. And that's the thing, is that people are not as familiar with the local large language model, partially because it's not kind of as front and center in the media and it's not as sexy sounding, but also because sometimes people actually in some websites are using them. They just don't realize that that's what they're using. So there are a lot of websites where the back end does not necessarily connect to Chat GPT or something like that, where they are running a local large language model on their servers, but um no one really is exposed to that unless they go through and read the fine print, which of course no one ever does. Got it. Sure, of course.

SPEAKER_00

Yeah. So walk me through maybe the nuts and bolts. If people are still listening, they're definitely intrigued, right? To this, most likely. Um I know that when Agentic AI pieces started coming out, you know, people were loading it and realizing, like, hey, to do things locally is pretty technical. Like, unless you're you know really in the industry, it can be kind of maybe daunting. Um what's it like to set up your app? Because you I assume there's some you know installation or or pieces required to get it running on whatever equipment you know they have.

SPEAKER_01

Well, so right now the version that I have is working, but it it's in a it's sort of like a working demo. So it really works, people can really use it, but it's a software as a service. So it's on a server that is not connected to Chat GPT and not connected to Cloud, and your information never leaves that one server. But if let's say a large company wanted to use this same technology and completely enclose it inside their walled garden, then they could easily just install all of the same software on their internal servers and use it there, or even on an individual computer and use it there. So what we have now is because it's just a lot in that case. Yeah. Yeah, exactly. They're licensing it. And so, you know, right now to demonstrate the technology, it's just a lot easier to have software as a service where you just go to a web page and do a few clicks rather than have someone download software and have to debug. I mean, there is actually a downloadable version of it, but we don't offer it right now because you know it's just there would be so much technical setup and it's a little bit of a good thing. Yeah. That that's why it's software software as a service now. But the way it works basically is you just go there, you open, you create an account, and we have free beta accounts available right now. And you can do a couple things. One is for free without even putting in, you know, not even having to open an account. You can get, you can have it analyze your voice. So you put in, paste in some examples of documents that are very characteristic of your writing. And it will give you a stylometric analysis that says things like you tend to use m-dashes more than most people, you tend to uh use abbreviations less than most people, you know, and so on and so forth. And what it also gives you for free is a prompt that you actually can use with ChatGPT or uh Claude. So you can just copy and paste this prompt and it will say, I speak like this, and it will give all the stylometrics of how the stylometric analysis has resolved. And it will give selected snippets of what you've said, not the whole thing because that would overload it, but just very characteristic snippets. So you can train other models better. Well, this is really not for training. This is for using um just as a sort of to give as an example of the kinds of things you can do when you personalize. You can use it with ChatGPT, you can use it with Cloud, but the results, as I've said, I mean, we've done comparisons, and it's better than if you didn't use it, but it's just still nowhere as good as the actual product, which is you go, you drag in. So you can actually, if you have a Substack, for example, you can just type in the address to your Substack and it will automatically crawl that or any RSS in your feed. Yeah. Yeah.

SPEAKER_00

Um, so if you have a blog or something already, you you can it'll just train correctly.

SPEAKER_01

Immediately do that. Right. And then if you have, let's say, all of your data are just in separate files on your machine, you can just upload that in one felt swoop, and then it goes to this protected server. And the server thinks about it for a while. It trains, it takes about 10, 20 minutes to train it, and then you can do a whole bunch of really cool stuff. Um, one of the things, the sort of basic first thing that you can do is you can give it some text and have it rewrite that in your voice. So let's say you have something that either you've written or someone else has written, but you're not sure that it's kind of on the brand voice, let's say that you're trying to recreate, or it's been written by ChatGPT, so it's very generic. Um, you just push a button and it'll rewrite it immediately and give you the same text, but in your voice. Or you can have it craft new documents in your voice. And this happens very quickly, you know, and it's basically very much like the normal Chat GPT experience, except it's doing something that's actually been trained on your real voice. And the really cool thing about it is that as you get more and more examples of your writing, so let's say you're writing a Substack and you have like 10 new articles every month or something like that, you can retrain it and have it update so that it reflects your most current kind of writing. Um, there's also a lot of really cool things you can do once you've modeled someone's voice in a neural network. Um, for example, you can take snapshots at different points in time and you can have an analysis that says, this is the way you were writing three months ago, and this is the way you're writing now, and there's been some drift, and the drift is in this direction. And that's just not something you could ever do with a frontier model. Um you can have it argue against itself. So it can take all of the information that it's pulled from what you've given it, and you can say, Hey, here's an idea for a new Substack article, let's say. Tell me if it contradicts anything that I've already said in any of my 100 or whatever Substack articles, or argue me out of saying it. And it'll be able to do all that in your voice. Um another feature partner, it's like it's a partner. You're helping you process that. Yeah. Helping you process that, but with your written words and with your written style. And it's very powerful. Um, another feature that I've been working on that's really, really cool, I think. I think this is kind of the killer app part of it, is a chat bot, but it's a chat bot that speaks in your voice. And you can even actually, if you clone your voice, you can actually even have it literally speak in your voice. And you interrogate the chat bot just as you would a normal chat bot about your work. So you can go in and you can say, Hey, what have I been talking about recently in my Substack? And what are some areas that maybe I've neglected that are still relevant to the field? And it will be able to look at everything that you've produced and give you, in your own words, a report. And you can say, Hey, that's an interesting topic. Write up that topic for me right now, and it'll do that.

SPEAKER_00

That's really cool. So you're yeah, let me put some examples out there. So maybe this would be maybe for publicly facing documents, what, emails, blog posts where you're generating a lot of content, you have a voice, but you don't want to manually do a lot of these things. So it's it's kind of morphing or combining the roles of maybe like a virtual assistant slash secretary with a copywriter. And the goal would be I mean, good grief, you could exponentially increase your efficiency, but still sound genuine. And it's almost like you're plagiarizing, but you're not, because you're giving it its own your voice and your content. I think that's exactly people crave that authenticity. You're doing that at scale with tools.

SPEAKER_01

Yes. So there are a lot of right. And the use cases are, for example, you know, you're a marketing agency and you have a whole bunch of brand documents that say this is how we want everything that goes out under our name to sound. And here are some examples of it. You can train it on that. Um, if you're onboarding new people, it becomes much easier because instead of having to have them go through some training course about the proper way to discuss the brand. What you can do is just have it use this tool and it will go ahead and produce everything in the correct brand voice. If you're selling the company, you can do the same thing. You can say, we have all these brand voice documents, but you don't even need to use those because you can just in day one type stuff into this interface and it will give you the kind of document that you're looking for. So it's got a lot of really exciting potential uses, I think.

SPEAKER_00

Yeah, I I completely agree. I mean I'm I'm exploding with ideas just talking to you about things with or any small business could really take advantage of this. What you mentioned this earlier a couple times, but I want to talk about the security aspect because I think that's really important. So you're basically giving people the benefit of a local LLM that's not tied to the big frontier models, but you're currently hosting it yourself, right? So it isn't technically local it's it's local to you and the beta user it's local to the server and it doesn't go out anywhere else.

SPEAKER_01

But there's nothing inherent at the technology that would keep someone you know you could do you could replicate all of that in a day trivially on let's say you're working for a big company and they want to use this or even a marketing company. You could replicate the entire thing easily well on under a day really because it's just copying files in order to have it literally air gapped. And so you would have basically a white labeled version of this that would only work inside your private network and you you know could conceivably you don't need to even be connected to the internet because it's never going out to the internet to do any sort of uh chat GPT or or Claude or any of that sort of stuff. Right. The only thing that you would really so if you're doing search obviously which you know you don't have to have but search would be something where you would still want to be connected to the internet obviously but the search doesn't send your information out onto the internet of course what it does is it just sends some you know keyword queries. But other than that, it's you know completely protected. And um so you know my view is the software as a service is probably good enough for a lot of people but if people are really super concerned about security then you know we can talk and just put it on there inside their network.

SPEAKER_00

Yeah. Well I would think you it would be something that people could host for themselves too. It's not like you have to have an office I mean you could throw up a an instance anywhere and and put it on a server. I think that's pretty common these days. Yeah where do you see this going, Arshavir? I mean I feel like you're just on the you're in beta testing now, right? So it's really early and you already have all these features and ideas. What's paint the vision in the next you know 12 to 18 months. What can people expect as maybe this becomes more commonplace and the I feel like we're on this AI curve, right? Like it's going to get worse before it gets better and you're you're in that better phase already, I feel like where do you see this going?

SPEAKER_01

Well I think the key to this particular technology to this product is that it's able to get better based on the input that you give it. So every time you enter in something in a chat window or every time you ask it a question, those are all potential things for it to learn about your personal style. So unlike with a Frontier model, which every time you go to it starts fresh, what you're able to do is build up a profile of your own work product and of your own content and interests and style. So style is the really you know the key thing here. It's not just content, but it's also style so that what comes out actually sounds like something you wrote. So I'm not sure where it's going to go um in the I mean, you know, the market kind of is what defines that. And so I have a lot of different ideas like some of the things I've already talked to you about. And I have idea I think that the chatbot thing is very powerful, but we'll have to see you know which aspects of this beta really appeal to people because you know I tend to take to software development as a very flexible process. So I don't really see creating a whole bunch of features and then you're just stuck with them. I have a actually on the website a list of what features would you like to see next and you can check off. You know, I think it would be interesting if it integrated with Evernote. I think it would be interesting if it integrated with Google things like that. So all of that is really kind of consequent to getting marketing feedback. And we're just very early in that stage. So you know come back to me in a few months and I can tell you uh hopefully that you know some people like this feature but I got rid of that feature because no one liked it. And you know I could be wrong about the chat bot. I mean it could be something that no one really finds particularly useful but I think it's going to be um a pretty cool thing because you get to basically talk to yourself. Yeah. And that's a very story.

SPEAKER_00

Yeah exactly but now it's you talk back right so um that's well you're unique in that you're um basically you're crowdsourcing almost a feature set and you're providing the tool and then your customers can take the tool and and enhance whatever they are doing, right? So your exact your um your reach and your ripple effect I think could be huge. Yeah.

SPEAKER_01

I mean I mean that's what I'm hoping. So you know we'll just have to see how the market reacts um as with anything and um you know we'll you know come back talk to me in three or six months and we'll see hopefully there will have been some forward motion in that in that time.

SPEAKER_00

Well I'll I'll definitely take you up on that I I keep a list of everyone on here because I I do want to follow up and continue the journey right it's not just a point in time but everybody's on their own path and I think that's exciting to see people develop. What would you tell people if they want to get a hold of you, Arshavir? I mean you are the company right now and we'll definitely put the website and everything in the show notes. In fact did we even say the name of the app it it's it's your voice your voicecraftai.

SPEAKER_01

And if you go to the terrible voicecraft towardsai that's okay. Yeah if you go to that you can do the free stylometric analysis and the free chat GPT prompt you don't have to put in any personal information. You just put in you know whatever text you wanted to learn. If you want to get a if you want to enroll in the beta trial there's a little beta trial uh button at the top and then we'll put you on the wait list and the beta trial will let you use all the features. So there's three tiers and the top tier is studio and that is the beta trial lets people use all the features of the product. And what's you know my goal or my hope is only that people doing a beta will give me lots and lots of feedback and tell me what they like, what they don't like what's broken, what's working. So that's the goal of the beta right now is to um really shake out all the bugs in it. It's exactly it's it's early software and so uh you know I'm sure there's a lot of places for improvement.

SPEAKER_00

Well you were you were kind enough to give me a beta a link you know earlier and I I was I was installing it and kicking the tires um and I even like I shared Bifi before I did give it to my CEO because we're struggling with um how to scale relationships right in in a world that's just overwhelmed with fake content and like people are craving the human touch. And yet that's really really hard to scale because you can only squeeze so much time out of a day. So yeah we'll we'll definitely be kicking the tires and I hope both our audience will check it out as well.

SPEAKER_01

Every day I add new features to it and there's a features list that I've put in there so you can know the chat bot I just added in very recently and I'm still adding features to the chat bot. And as I said I my intuition is that's going to be the killer aspect of this app is that to be able to interact with your own written material in your voice and to have it respond in your voice and then when it produces new product have that product have that written product be in your voice as well. Yeah yeah I love that.

SPEAKER_00

Well I've really enjoyed talking to you today. Thank you so much for your time um like I said we'll we'll put all the links in our our show notes so that people can reach out to you and and hopefully get into this product more. And um yeah we'll definitely follow up in a few months and and continue to follow along the journey. I think there's a ton of potential here and um hopefully people can check it out and it it can help them grow and get better in their their own roles. So I hope so thank you so much. Yeah thank you I enjoyed it. You bet. Take care