<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
<style type="text/css" style="display:none;"> P {margin-top:0;margin-bottom:0;} </style>
</head>
<body dir="ltr">
<div style="font-family: Aptos, Aptos_EmbeddedFont, Aptos_MSFontService, Calibri, Helvetica, sans-serif; font-size: 12pt; color: rgb(0, 0, 0);" class="elementToProof">
OK, I had thought perhaps this was local storage, but it's networked, understood.</div>
<div style="font-family: Aptos, Aptos_EmbeddedFont, Aptos_MSFontService, Calibri, Helvetica, sans-serif; font-size: 12pt; color: rgb(0, 0, 0);" class="elementToProof">
<br>
</div>
<div style="font-family: Aptos, Aptos_EmbeddedFont, Aptos_MSFontService, Calibri, Helvetica, sans-serif; font-size: 12pt; color: rgb(0, 0, 0);" class="elementToProof">
But <b>why</b>&nbsp;are we running s3fs or gcfuse on it?&nbsp; What's the purpose/benefit of doing this?&nbsp; The stored data is Lustre stripe objects, not whole files, so not as a practical matter useful without Lustre.&nbsp; So why not format it as an OST and use it directly?</div>
<div style="font-family: Aptos, Aptos_EmbeddedFont, Aptos_MSFontService, Calibri, Helvetica, sans-serif; font-size: 12pt; color: rgb(0, 0, 0);" class="elementToProof">
<br>
</div>
<div style="font-family: Aptos, Aptos_EmbeddedFont, Aptos_MSFontService, Calibri, Helvetica, sans-serif; font-size: 12pt; color: rgb(0, 0, 0);" class="elementToProof">
-Patrick</div>
<div id="appendonsend"></div>
<hr style="display:inline-block;width:98%" tabindex="-1">
<div id="divRplyFwdMsg" dir="ltr"><font face="Calibri, sans-serif" style="font-size:11pt" color="#000000"><b>From:</b> Jinshan Xiong &lt;jinshanx@google.com&gt;<br>
<b>Sent:</b> Tuesday, November 4, 2025 1:02 PM<br>
<b>To:</b> Patrick Farrell &lt;pfarrell@ddn.com&gt;<br>
<b>Cc:</b> Oleg Drokin via lustre-devel &lt;lustre-devel@lists.lustre.org&gt;; Oleg Drokin &lt;green@whamcloud.com&gt;<br>
<b>Subject:</b> Re: [lustre-devel] RFC: Spill device for Lustre OSD</font>
<div>&nbsp;</div>
</div>
<div>
<div dir="ltr">
<div dir="ltr"><br>
</div>
<br>
<div class="x_gmail_quote x_gmail_quote_container">
<div dir="ltr" class="x_gmail_attr">On Tue, Nov 4, 2025 at 10:39\u202fAM Patrick Farrell &lt;<a href="mailto:pfarrell@ddn.com">pfarrell@ddn.com</a>&gt; wrote:<br>
</div>
<blockquote class="x_gmail_quote" style="margin:0px 0px 0px 0.8ex; border-left:1px solid rgb(204,204,204); padding-left:1ex">
<div class="x_msg-774908302111501245">
<div dir="ltr">
<div style="font-family:Aptos,Aptos_EmbeddedFont,Aptos_MSFontService,Calibri,Helvetica,sans-serif; font-size:12pt; color:rgb(0,0,0)">
Right, but if this spill device is separated from the OST, how can it fail over with the OST?&nbsp; If it
<i>can</i>&nbsp;fail over with the OST, like eg it's just another network attached device that the OST would access somehow*, why not either use it for the principal OST or make it a separate OST?</div>
</div>
</div>
</blockquote>
<div><br>
</div>
<div>I'm not sure I understand the question. I don't know why failover would get in the way.&nbsp;It would just umount primary OSD and then spill device, and then reverse the process on the new OSS.</div>
<div><br>
</div>
<div>An OST requires a fully&nbsp;functional OSD. The spill device would use s3fs or gcsfuse, so it won't be a fully functional OSD.</div>
<div>&nbsp;</div>
<blockquote class="x_gmail_quote" style="margin:0px 0px 0px 0.8ex; border-left:1px solid rgb(204,204,204); padding-left:1ex">
<div class="x_msg-774908302111501245">
<div dir="ltr">
<div style="font-family:Aptos,Aptos_EmbeddedFont,Aptos_MSFontService,Calibri,Helvetica,sans-serif; font-size:12pt; color:rgb(0,0,0)">
<br>
*this is actually ignoring several questions about failover to another machine - for example, unless you plan to have a clustered file system on it, writing to it from a different location requires a handoff and exclusion, etc<br>
<br>
</div>
<div style="font-family:Aptos,Aptos_EmbeddedFont,Aptos_MSFontService,Calibri,Helvetica,sans-serif; font-size:12pt; color:rgb(0,0,0)">
-Patrick</div>
<div id="x_m_-774908302111501245appendonsend"></div>
<hr style="display:inline-block; width:98%">
<div id="x_m_-774908302111501245divRplyFwdMsg" dir="ltr"><font face="Calibri, sans-serif" color="#000000" style="font-size:11pt"><b>From:</b> Jinshan Xiong &lt;<a href="mailto:jinshanx@google.com" target="_blank">jinshanx@google.com</a>&gt;<br>
<b>Sent:</b> Tuesday, November 4, 2025 11:38 AM<br>
<b>To:</b> Patrick Farrell &lt;<a href="mailto:pfarrell@ddn.com" target="_blank">pfarrell@ddn.com</a>&gt;<br>
<b>Cc:</b> Oleg Drokin via lustre-devel &lt;<a href="mailto:lustre-devel@lists.lustre.org" target="_blank">lustre-devel@lists.lustre.org</a>&gt;; Oleg Drokin &lt;<a href="mailto:green@whamcloud.com" target="_blank">green@whamcloud.com</a>&gt;<br>
<b>Subject:</b> Re: [lustre-devel] RFC: Spill device for Lustre OSD</font>
<div>&nbsp;</div>
</div>
<div>
<div dir="ltr">
<div dir="ltr"><br>
</div>
<br>
<div>
<div dir="ltr">On Mon, Nov 3, 2025 at 6:47\u202fPM Patrick Farrell &lt;<a href="mailto:pfarrell@ddn.com" target="_blank">pfarrell@ddn.com</a>&gt; wrote:<br>
</div>
<blockquote style="margin:0px 0px 0px 0.8ex; border-left:1px solid rgb(204,204,204); padding-left:1ex">
<div>
<div><span>I haven\u2019t seen any mention of failover yet in this conversation (may have missed it), but if the device is truly local, then in failed over configurations the data is inaccessible.&nbsp; If it\u2019s *not* local, why not just make the device part of the OST
 or an independent OST?</span></div>
</div>
</blockquote>
<div><br>
</div>
<div>It won't be local. Actually, this is designed for the cloud.</div>
<div><br>
</div>
<div>We already have tiered storage based on mirroring; however, that still requires clients to move data and a file system level scanner to decide which files move to the cold tier. It's cumbersome&nbsp;to maintain those clients.</div>
<div><br>
</div>
<div>&nbsp;</div>
<blockquote style="margin:0px 0px 0px 0.8ex; border-left:1px solid rgb(204,204,204); padding-left:1ex">
<div>
<div>&nbsp;</div>
</div>
</blockquote>
<blockquote style="margin:0px 0px 0px 0.8ex; border-left:1px solid rgb(204,204,204); padding-left:1ex">
<div>
<div><br>
</div>
<div><br>
</div>
<div>Or even if it is local only - it could be made an OST</div>
<div></div>
<hr style="display:inline-block; width:98%">
<div id="x_m_-774908302111501245x_m_-3839173029790542839divRplyFwdMsg" dir="ltr">
<font face="Calibri, sans-serif" color="#000000" style="font-size:11pt"><b>From:</b> lustre-devel &lt;<a href="mailto:lustre-devel-bounces@lists.lustre.org" target="_blank">lustre-devel-bounces@lists.lustre.org</a>&gt; on behalf of Oleg Drokin via lustre-devel &lt;<a href="mailto:lustre-devel@lists.lustre.org" target="_blank">lustre-devel@lists.lustre.org</a>&gt;<br>
<b>Sent:</b> Monday, November 3, 2025 7:58 PM<br>
<b>To:</b> <a href="mailto:jinshanx@google.com" target="_blank">jinshanx@google.com</a> &lt;<a href="mailto:jinshanx@google.com" target="_blank">jinshanx@google.com</a>&gt;<br>
<b>Cc:</b> <a href="mailto:lustre-devel@lists.lustre.org" target="_blank">lustre-devel@lists.lustre.org</a> &lt;<a href="mailto:lustre-devel@lists.lustre.org" target="_blank">lustre-devel@lists.lustre.org</a>&gt;<br>
<b>Subject:</b> Re: [lustre-devel] RFC: Spill device for Lustre OSD</font>
<div>&nbsp;</div>
</div>
<div><font size="2"><span style="font-size:11pt">
<div>On Mon, 2025-11-03 at 16:33 -0800, Jinshan Xiong wrote:<br>
&gt; &gt; &gt; &gt; <br>
&gt; &gt; &gt; &gt; I am not sure it's a much better idea than the already existing<br>
&gt; &gt; &gt; &gt; HSM<br>
&gt; &gt; &gt; &gt; capabilities we have that would allow you to have &quot;offline&quot;<br>
&gt; &gt; &gt; &gt; objects<br>
&gt; &gt; &gt; &gt; that would be pulled back in when used, but are otherwise just<br>
&gt; &gt; &gt; &gt; visible<br>
&gt; &gt; &gt; &gt; in the metadata only.<br>
&gt; &gt; &gt; &gt; The underlying capabilities are pretty rich esp. if we also<br>
&gt; &gt; &gt; &gt; take<br>
&gt; &gt; &gt; &gt; into<br>
&gt; &gt; &gt; &gt; account the eventual WBC stuff.<br>
&gt; &gt; &gt; <br>
&gt; &gt; &gt; The major problem of current HSM is that it has to have dedicated<br>
&gt; &gt; &gt; clients to move data. Also, scanning the entire Lustre file<br>
&gt; &gt; &gt; system<br>
&gt; &gt; <br>
&gt; &gt; This (dedicated client) is an implementation detail. It could be<br>
&gt; &gt; improved in many ways and the effort spent on this would bring<br>
&gt; &gt; great<br>
&gt; &gt; benefit to everyone?<br>
&gt; <br>
&gt; <br>
&gt; Almost all designs assume some upfront implementation. We (as the<br>
&gt; Lustre team) considered running clients on OST nodes, but cloud users<br>
&gt; are sensitive about their data being exposed elsewhere.<br>
&gt; <br>
&gt; Can you list a few improvements that come to mind?<br>
<br>
Clients between server nodes is probably one of the most obvious<br>
choices indeed. Considering the data already resides on those nodes, I<br>
am not sure I understand the concerns about &quot;exposing&quot; data that's<br>
already on those nodes. If customers are so sensitive, we support data<br>
encryption.<br>
We could also do some direct server-server migration of some sort where<br>
OSTs exchange data without bringing up real clients and doing copies<br>
from userspace. That might be desirable for other reasons for future<br>
functionality (e.g. various caching things people have been long<br>
envisioning)<br>
&nbsp;<br>
&gt; &gt; <br>
&gt; &gt; &gt; takes very long time so it resorts to databases in order to make<br>
&gt; &gt; &gt; correct decisions about which file should be released. By the<br>
&gt; &gt; &gt; time,<br>
&gt; &gt; &gt; the two system will be out of sync. That makes it practically<br>
&gt; &gt; &gt; unusable.<br>
&gt; &gt; <br>
&gt; &gt; This again is an implementation detail, not even hardcoded<br>
&gt; &gt; anywhere.<br>
&gt; &gt; How do you plan for the OST to to know what stuff is not used<br>
&gt; &gt; without<br>
&gt; &gt; resorting to some database or scan? Now take this method and make<br>
&gt; &gt; it<br>
&gt; &gt; report &quot;upstream&quot; where currently HSM implementations resort to<br>
&gt; &gt; databases or scans.<br>
&gt; &gt; <br>
&gt; <br>
&gt; The assumption is that OST sizes are relatively small, up to 100TB.<br>
&gt; Also, scanning local devices in kernel spaces is much faster. So yeah<br>
&gt; there is no database in the way.<br>
<br>
I am not sure why? In general small OSTs are a relatively rare thing<br>
because to reach large FS sizes you would need to many of them, space<br>
balancing becomes a chore and so on. So relatively few do it for some<br>
fringe reasons (e.g. Google) . Majority of people prefer large OSTs.<br>
<br>
Also nothing stops you from doing a per-OST scan (when you do have<br>
small OSTs) and then kicking the results up to the acting agent to do<br>
something about it (or the other way around, the hsm engine can ask<br>
OSTs one by one (picking less busy ones or ones that have the least<br>
free space, or some other factor). And there's absolutely no need to<br>
wait out to query all OSTs, you can get results from one and work on<br>
the data from it while the other OSTs are still thinking (or not,<br>
there's absolutely no requirement to get full filesystem data before<br>
making any decisions).<br>
<br>
&gt; I guess users won't have 1PB OSTs, will they?<br>
<br>
There probably are already? NASA has a known 0.5P OST configuration:<br>
<a href="https://www.nas.nasa.gov/hecc/support/kb/lustre-progressive-file-layout-(pfl)-with-ssd-and-hdd-pools_680.html#:~:text=The%20available%20SSD%20space%20in%20each%20filesystem,decimal%20(far%20right)%20labels%20of%20each%20OST" target="_blank">https://www.nas.nasa.gov/hecc/support/kb/lustre-progressive-file-layout-(pfl)-with-ssd-and-hdd-pools_680.html#:~:text=The%20available%20SSD%20space%20in%20each%20filesystem,decimal%20(far%20right)%20labels%20of%20each%20OST</a><br>
.<br>
<br>
&gt; &gt; Rereading your proposal, I see that this particular detail is not<br>
&gt; &gt; covered and it's just assumed that &quot;infrequently accessed data&quot;<br>
&gt; &gt; would<br>
&gt; &gt; be somehow known.<br>
&gt; <br>
&gt; I should have mentioned that in the migration section. Also, we need<br>
&gt; to slightly update the OST read to use a local transaction to update<br>
&gt; an object's access time (atime) if it's older than a predefined<br>
&gt; threshold, for example, 10 minutes.&nbsp;<br>
<br>
This is going to be fragile in the face of varying clock times on<br>
different clients potentially not synced with the servers.<br>
Also in the face of -o noatime.<br>
<br>
But yes, I guess it's one way to get this &quot;on the cheap&quot;, and the other<br>
trouble I foresee is you are going to have a biased set. Only recently<br>
touched objects (so with fresh atime), unless you plan to retain a<br>
database and update it from such transaction flow, which certainly is<br>
possible, but I am not sure how practical vs some sort of a scan.<br>
<br>
&gt; &gt; &gt; &gt; If the argument is &quot;but OSTs know best what stuff is used&quot;<br>
&gt; &gt; &gt; &gt; (which I<br>
&gt; &gt; &gt; &gt; am<br>
&gt; &gt; &gt; &gt; not sure I buy, after all before you could use something off<br>
&gt; &gt; &gt; &gt; OSTs<br>
&gt; &gt; &gt; &gt; you<br>
&gt; &gt; &gt; &gt; need to open a file I would hope) even then OSTs could just<br>
&gt; &gt; &gt; &gt; signal<br>
&gt; &gt; &gt; &gt; a<br>
&gt; &gt; &gt; &gt; list of &quot;inactive objects&quot; that then a higher level system<br>
&gt; &gt; &gt; &gt; would<br>
&gt; &gt; &gt; &gt; take<br>
&gt; &gt; &gt; &gt; care of by relocatiing somewhere more sensical and changing the<br>
&gt; &gt; &gt; &gt; layout<br>
&gt; &gt; &gt; &gt; to indicate those objects now live elsewhere.<br>
&gt; &gt; &gt; &gt; <br>
&gt; &gt; &gt; &gt; The plus here is you don't need to attach this &quot;Wart&quot; to every<br>
&gt; &gt; &gt; &gt; OST<br>
&gt; &gt; &gt; &gt; and<br>
&gt; &gt; &gt; &gt; configure it everywhere and such, but rather have a central<br>
&gt; &gt; &gt; &gt; location<br>
&gt; &gt; &gt; &gt; that is centrally managed.<br>
&gt; &gt; &gt; &gt; <br>
&gt; &gt; <br>
&gt; &gt; _______________________________________________<br>
&gt; &gt; lustre-devel mailing list<br>
&gt; &gt; <a href="mailto:lustre-devel@lists.lustre.org" target="_blank">lustre-devel@lists.lustre.org</a><br>
&gt; &gt; <a href="http://lists.lustre.org/listinfo.cgi/lustre-devel-lustre.org" target="_blank">
http://lists.lustre.org/listinfo.cgi/lustre-devel-lustre.org</a><br>
&gt; &gt; <br>
<br>
_______________________________________________<br>
lustre-devel mailing list<br>
<a href="mailto:lustre-devel@lists.lustre.org" target="_blank">lustre-devel@lists.lustre.org</a><br>
<a href="http://lists.lustre.org/listinfo.cgi/lustre-devel-lustre.org" target="_blank">http://lists.lustre.org/listinfo.cgi/lustre-devel-lustre.org</a><br>
</div>
</span></font></div>
</div>
</blockquote>
</div>
</div>
</div>
</div>
</div>
</blockquote>
</div>
</div>
</div>
</body>
</html>