What if the hardest problem in AI networking isn't speed, but keeping every link running?
In this episode, CEO Martin Lund argues that as AI outgrows a single rack, the interconnect becomes part of the computer itself, and reliability becomes the requirement that cannot be traded away.
Martin started in networking on dial-up modems, and he calls AI scale-up "the most unforgiving, demanding traffic you can put on a network": latency intolerant, loss intolerant, and growing with every generation. His answer is to optimize for the whole AI factory instead of a single component, and to reject fixes that only look good locally, like replaceable lasers that still need a person to swap them. As Martin puts it, a $50 billion data center should not run at 50% utilization because the technician on laser duty was out sick. CScale integrates the laser so failures are handled inside the system, building toward links that need zero touch.
In this conversation, we cover:
- Why CScale optimizes for "the factory," not "the individual rack"
- Why "serviceability is not a solution for continuity" in AI data centers
- How scale-up differs from scale-out, and why tightly coupled accelerators are "actually the machine"
- What "the copper wall" is, and why optics is the way past it
- How designing around the fact that "lasers will fail" lets the scale-up network "behave like copper"
Martin also explains why fewer fibers mean "fewer things to go wrong," why CScale is not "building a company to solve one generation," how the team blends deep domain experts in a collaborative culture, and how better factory utilization can lower the power consumption of AI.
CScale is an optical interconnect company building an integrated light engine for AI scale-up, designed to contain optical failures so thousands of accelerators can keep working as one much larger computer.
[00:00:00] Martin: The factory we optimize it for. It's not the individual rack
[00:00:12] Martin: I'm Martin Land. I am the CEO of CScale, and we are building optical interconnect for scale-up solutions for the future.
[00:00:19] Sandesh: Hi, I'm
[00:00:19] Sandesh: Sandesh Patnam.
[00:00:20] Sandesh: We're excited to announce, uh, today that we're co-leading an investment in CScale, $145 million round, alongsideSutter Hill Ventures and Atreides Capital. At a high level, it's a actually a pretty simple secular trend that we are leaning into.
[00:00:37] Sandesh: As we know, AI requires a lot of compute and a lot of data that needs to be exchanged, and the requirements that it places on it- on networking is dramatically different. To date, copper has been the default solution, and we all know that over time optics is inevitable, and that's the fundamental problem that CScale is here to solve.
[00:00:58] Sandesh: Martin Land, CEO of CScale, joining us today as part of our CEO series. Welcome, Martin. Thank you very much. Your background in networking goes a long way, and there are lots and lots of different generations of technology you have seen come and go. What's unique about, you know, this specific iteration, especially with AI?
[00:01:16] Martin: It's true. I've, I've been around networking a bit. I've, uh, I like to say I started with the dial-up modems and, and, and modems running at 300 baud. So I've seen a lot of changes and, and, and speed upgrades, uh, throughout my career. Uh, what's unique about, uh, AI networking or what we talk about scale-up networking, is that it is the most unforgiving, demanding n- traffic you can put on a network.
[00:01:44] Martin: It is latency intolerant. You need to be deterministic. It is loss intolerant, means it doesn't really like that you drop packets. The bandwidth requirements are immense and, and continuously increasing. Every generation, it doubles, quadruples maybe, and you have to manage the power and the complexity around operational, having all these links i-i-in a network.
[00:02:09] Martin: So it basically takes all the problems in networking that we have sort of solved individually over time, put them all together, and then put it on a, on a roadmap of change that it, it's like a double every two year. The interconnect for, for AI, it's really a compute fabric. It's not really networking as much as it's like interconnecting these accelerators is what is governing the network scalability and the economic scalability, and, uh, so it's super, super critical.
[00:02:39] Martin: It's a... It's, it's clearly the next frontier.
[00:02:41] Sandesh: Yeah. You mentioned some fascinating things here. One is a systems approach, right? You have to have a full systems approach. Our view is in a lot of the optical startups that are trying to address this, it's typically a single component, something that's about speed or maybe it's the laser that, that is unique, and everything else that they do is in benefit of that specific innovation, so a very sort of component-driven approach to solving the problem.
[00:03:11] Sandesh: Something unique about C-Scale that really was impressive is, when we started spending time with you all, was how you went about it from a first-principle thinking. Almost no priors, go back to basics, and rethink the problem from a systems perspective. Tell us more about that approach to the problem that's unique and differentiated a little bit to everybody else.
[00:03:31] Martin: Yeah, I mean, it's, it's the old saying is, "Everything you have is a hammer, everything looks like a nail," and, and, and AI is a systems optimization problem. It, it's really about not only can you make a better link, uh, or, or make it faster or, or lower power or, or better reliable, it's all of the above, but actually it's about the, as Jensen calls it, the AI factory.
[00:03:58] Martin: It's the factory we optimizing for. Mm-hmm. It's not the individual rack. And everything that we do in C-Scale is centered around that principle, that we're optimizing for the s- system-
[00:04:08] Sandesh: Right ...
[00:04:08] Martin: the solution, not the individual.
[00:04:09] Sandesh: Right.
[00:04:10] Martin: I- we can give one example.
[00:04:11] Sandesh: Yeah.
[00:04:12] Martin: You can, uh, come up with a solution that, uh, uh, which is very efficient and very good and, but it, it may, let's say, it requires you to have 100 times more fiber infrastructure Than if you, let's say, go with what we have, right?
[00:04:30] Martin: So 100X problem gets exported, but it's not y- your little device may be cheaper and, and, and it performs very well, but the problem is now for the operator that's gonna install the, the two orders of magnitude more cables and infrastructure, and the cost associated with, that may not show up.
[00:04:49] Sandesh: Right.
[00:04:50] Martin: And what's the cost of replacement, and what's the failure rate on that?
[00:04:53] Martin: Or another example is you can have maybe part of the system that is replaceable. Mm-hmm. Like, some people will have replaceable lasers. Like, lasers fail. They will fail, so let's make them replaceable. But locally, it seems like a good idea.
[00:05:07] Sandesh: Right.
[00:05:08] Martin: We have a point of view that says, no, it's not a good idea.
[00:05:11] Martin: Serviceability is not a solution for continuity.
[00:05:15] Sandesh: Right.
[00:05:15] Martin: It's a different approach. Like, a $50 billion data center, you don't want it to run it at 50% utilization because Joe, uh, that is doing re- laser replacement, uh, was out sick. Right. Right.
[00:05:29] Sandesh: Right.
[00:05:29] Martin: And there's... Like, it's just not practical. Right. Uh, so if you think of all these systems, elements, and problems that come together, and we kind of boil it down to what you're trying to do, is to build something that's really good Really efficient and very simple.
[00:05:47] Martin: Actually, you don't wanna touch it. Zero touch is one of the things we say we are focusing on
[00:05:51] Sandesh: Some very fascinating points that you raised there. Perhaps maybe let's go down the path of scale-up versus scale-out.
[00:05:57] Martin: Yeah.
[00:05:58] Sandesh: So as people who, who spend time in the scale-out domain, we already see optics today in the scale-out domain, different types of technologies, different trade-offs.
[00:06:05] Sandesh: And now the systems approach you just talked about, uniquely for scale-up, creates a whole bunch of different set of trade-offs, one of which was reliability, and the second is just in terms of, you know, the compute farm, right? You know, you have cascading failure if you have a specific link go down. So the industry is trying to solve all of these things with various approaches, and perhaps you can spend a little bit more time, but talk about something specific that scale-up presents that perhaps scale-out doesn't.
[00:06:34] Martin: The scale-up is different. This is where the, the accelerators gets connected.
[00:06:38] Sandesh: Mm-hmm.
[00:06:39] Martin: And, and each accelerator, uh, like if you look at how much bandwidth that wants, it's, it's actually... 'Cause you wanna talk to another accelerator, actually wanna talk to- Many ... many accelerators. Yeah. But the bandwidth between them is, it has to be non-blocking, and it has to be predictable.
[00:06:58] Martin: When you put them together, that's actually the machine.
[00:07:00] Sandesh: Right.
[00:07:01] Martin: It is not, they're not, they're not loosely coupled, they're tightly coupled. So what that does is it puts requirements on the, the, the interconnect, uh, to be super high bandwidth, super low latency. There's no time for going and replacing modules- Right
[00:07:17] Martin: because these devices are also way more expensive- Right ... than your traditional server. Right. So a long-winded way to say it's different. Bandwidth is probably orders of magnitudes more per device, plus the cost of downtime is order of magnitude higher. Uh, so it's a different, it's a wh- completely different game.
[00:07:37] Sandesh: Yeah, I think, so that raises the bar a lot for scale-up, right? So the trade-offs are very different, and people have talked about this copper wall, right? So today, the industry is basically saying, "Look, I'm gonna take copper compute to compute, scale up, and to the point where, you know, I'm gonna sweat copper to the very end."
[00:07:56] Sandesh: And then I'm gonna switch over to optics, right? So the bar is, you know, significantly higher as you, as you mentioned. Second thing that I think you talked about is sort of this notion of uptime and reliability being such a critical component of it, and then there are some industry movements in this direction, one of which is sort of, if you go down a technical path, is this notion of, you know, the laser itself, right?
[00:08:18] Sandesh: One is, you know, the power of the laser, the uptime of the laser, laser getting disaggregated from the actual connection because you don't wanna create the same level of heat dissipation, and then you are taking a slightly different approach. You're saying, "Let's bring the laser back on." Something unique about your insights allowed you to go down this path.
[00:08:38] Sandesh: Can you comment about that and tell us what was the specifics, insights that led you to make what I would say is a contrary decision today to take lasers on,
[00:08:46] Martin: on? Yeah, it is. It is. And, and we can start with, with the laser pieces to say- Mm-hmm ... um, i- if it's not integrated, you can't really manage it and you can't do the co-optimization with it.
[00:08:57] Martin: Right. That's one piece.
[00:08:58] Sandesh: Right.
[00:08:58] Martin: We know lasers will fail.
[00:09:00] Sandesh: Right.
[00:09:01] Martin: Let's just accept that. Let's just come up with a system that- Right ... that can, can, can handle that.
[00:09:06] Sandesh: Right.
[00:09:07] Martin: And this is some of the unique stuff that we're doing. We're saying, "Lasers will fail, and we manage that. The system outside don't need to worry about it."
[00:09:15] Sandesh: Mm-hmm. "
[00:09:15] Martin: We will deal with it." So it's invisible.
[00:09:17] Sandesh: Right.
[00:09:18] Martin: Laser fails, others, the system won't know.
[00:09:21] Sandesh: Right.
[00:09:21] Martin: Great. The scale-up network wants to behave like copper.
[00:09:25] Sandesh: Mm-hmm.
[00:09:25] Martin: And maybe you think of copper, we're talking about the copper wall. What is this copper wall? Right. What is it like? Right. Copper is a great, uh, it's a great thing.
[00:09:33] Martin: Inside a chip, it's, there's like copper wires- Yep, right ... that connect p- different part of the chip together. Right. Nobody worries about that. It just works- Right ... all the time. And on the PCB, you have a, you interconnect compute devices together. That also works all the time, and it does that because distance are short.
[00:09:49] Sandesh: Right.
[00:09:49] Martin: And the faster your signal runs, the, the harder it is to get them across, uh, to a longer distance. Therefore, when you then wanna put more compute into a rack, you, and you still wanna connect it with copper 'cause it behaves so well, it's highly reliable, it's predictable, I need to use maybe a copper back plane- Mm-hmm
[00:10:10] Martin: uh, or, uh, fly over cables. All this stuff is done to keep that going.
[00:10:14] Sandesh: Right.
[00:10:15] Martin: But now we've, now we're running out of space in that rack.
[00:10:17] Sandesh: Right.
[00:10:18] Martin: And even from a power cooling perspective, it, it's very clear that it's hard to put more into the rack, more power into the rack and so forth, so it becomes like maybe I can put 72 of these devices into a rack of these accelerators.
[00:10:33] Martin: I have to put, break them into more racks. Once you get out of that, the distance, and this is the copper wall- We're talking about is like, it may be okay to go one to two meters- Right ... when you're running at these speeds. It's not, you can't really go 50, 100 meters on this, or 200 meters, 500 meters. You can't.
[00:10:53] Martin: You have to move into another medium, which is optics is the way to do it. Optics, yeah. Our unique insight is to say, "So okay, that's gonna happen."
[00:10:59] Sandesh: Right.
[00:10:59] Martin: But we don't wanna have optics, which usually is, it, it's fragile.
[00:11:03] Sandesh: Right.
[00:11:04] Martin: Compared to copper, it's fragile. Compared, yeah. In the scale-up domain, you don't really have that luxury.
[00:11:08] Martin: You have to have, uh, very high reliable, uh, links, uh-
[00:11:12] Sandesh: Right ...
[00:11:13] Martin: maybe even more than in any other- Right ... system in order to keep the traffic flowing.
[00:11:18] Sandesh: You know, any insights on s- you know, how people in industry are thinking about spectrum, and is there anything unique about C Scale and how we've thought about spectrum?
[00:11:26] Martin: Well, I think we're, we definitely believe that it's good to have fewer fibers.
[00:11:32] Sandesh: Right.
[00:11:32] Martin: You know, fewer physical things means fewer phi- things to go wrong. What we believe is that once we're in the light domain, let's use the fact that we're in the, the light domain, and we can put a lot more bandwidth down a fiber-
[00:11:50] Sandesh: Right
[00:11:50] Martin: if, if we do it right than, than, uh, if we do it wrong. We're not building a company to solve one generation, right? We believe AI scale-up will continue to demand more and more and more and more bandwidth. There's the treadmill.
[00:12:07] Sandesh: Right.
[00:12:08] Martin: I don't see an end in sight on this. Right. So what means, it means that whatever we do has to have five, six generations of runway, that we have reasonable confidence that we can get to.
[00:12:18] Sandesh: Right.
[00:12:19] Martin: The path we pick allows us to scale-
[00:12:22] Sandesh: Right ...
[00:12:23] Martin: versus we run into ro- a bottleneck that's hard, a bottleneck maybe two years from now, we just can't put more fibers into the box.
[00:12:31] Sandesh: Right,
[00:12:32] Martin: right. We run out of, out of space.
[00:12:33] Sandesh: Right. And
[00:12:33] Martin: I think that's, that's a unique insight in, in what we're doing in the spectrum.
[00:12:37] Martin: Uh, having a plenty of spectrum to play with is key. By the, it's interesting, Ron, it's actually what we're running out of- Right ... copper, is, is bandwidth.
[00:12:46] Sandesh: Right. Such a big aspect of what you're doing is, you know, bringing people from various domains and also, you know, to collaborate with each other to build something that's unique, uh, for this industry.
[00:12:57] Sandesh: Tell, tell us more about, you know, the unique team build, uh, that you're embarking upon here.
[00:13:02] Martin: I mean, first of all, we have a, a very, very talented team. Uh, incredible people in, in, in each of their domains, and the people we wanna bring in are really w- ones that are, that, that wanna be part of, of building something exciting, that are willing, uh, interested in working in a collaborative culture.
[00:13:21] Martin: It's not about sitting over in your, in your cube and, and, and, and work by yourself. This is truly about collaboration because it's a sy- it's a systems optimization problem. We like to have people that wanna do hard things well- And want to collaborate with others in, in doing so. And, and we have lots of, of, of, of open positions, so we're blending experienced talent with, with, with deep expertise in certain domains and, and, and people that have lots of energy, uh, uh, for, for going and tackling, uh, hard problems, and that's what we're all about.
[00:13:58] Sandesh: The scale-up is the next sort of big sort of wall that one has to cross. Yeah. You solve an incredible problem. It requires a systems approach. You're building a absolutely phenomenal team, attracting some great talent, can work as a team. And then the impact, you know, that if we are successful, that it can have broadly on the AI, uh, adoption is gonna be massive, you know?
[00:14:22] Sandesh: So if you think about just sort of pure, like, the longevity of the value proposition that you bring to bear to the market seems enormous.
[00:14:30] Martin: I believe so, and I, I mean-
[00:14:32] Sandesh: Yeah ...
[00:14:33] Martin: we do a good job here. We'll also lower the power consumption of AI. And a GPU that's idling because waiting for traffic is- Right
[00:14:40] Martin: burning roughly the same amount of power as one that's doing good productive work.
[00:14:43] Sandesh: Right.
[00:14:44] Martin: Uh, a- and, and if we can improve the factory utilization by 10%-
[00:14:49] Sandesh: Right ...
[00:14:49] Martin: you know, uh, that's meaningful-
[00:14:50] Sandesh: Meaningful ...
[00:14:51] Martin: at scale. Better, faster, cheaper, more efficient communication is never go out of sp- uh, style.
[00:14:56] Sandesh: We are quite excited about what you guys are building.
[00:14:59] Sandesh: We think that, you know, from everything that we have seen, that I think I would say the chances of success with us is, is greater than anything that we have seen. Uh, we're here to support you on your journey, and, uh, excited to see what you're, what you're building and what's next.
[00:15:13] Martin: I'm very excited for your support.
[00:15:15] Martin: Thank you so much. And- Yeah ... and for the vision that you have.

